GPT-6 Luna feels like a strange trade-off
The early results on GPT-6 Luna make it look like a very strange model. OpenAI is presenting Luna as a cheaper, more efficient successor to GPT-5.6 Luna, while Artificial Analysis reports a mixed picture: some agent and terminal evaluations improve, but Luna regresses on GDPval and Briefcase. That fits my own experience. Luna feels much chattier than before, but I do not feel a corresponding increase in intelligence. The engineering has probably been streamlined quite heavily, which makes me wonder whether we are looking at a smaller model that has been pre-trained and post-trained to stay close to GPT-5.6 Luna on the benchmarks that matter most to the launch.
My guess is that OpenAI is aiming for a lower price point and has made Luna comparable to its predecessor across much of the headline testing, while accepting losses elsewhere. I will wait another week before forming a firmer view. The model is new, and OpenAI will keep post-training it, so the next set of updates may tell us whether this is a stable trade-off or just an awkward first release. (OpenAI’s launch report, Artificial Analysis’ evaluation)