Reflection’s Beam activates 23 billion of 501 billion parameters per token and, according to the company, matches GLM 5.2 on demanding reasoning tasks with three to four times less compute. Its reinforcement-learning phase used 10,500 Nvidia GB300 GPUs for over four weeks.