Artificial Analysis
- Indexed articles, last 90 days
- 4
- Latest publication
- Sep 29, 2026
- Outlet visibility, for Artificialanalysis
- Top 100K sites
- Earliest in this view
- Jul 21, 2026
Latest articles
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence (opens the original)
Read excerpt
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task. Pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, except that the cache read discount rises from 90% to 95%. GPT-6.1 Sol’s overall blended price for agentic workloads is therefore slightly lower than GPT-6 Sol. This represents an additional price cut, following GPT-6 Sol’s original 50% discount from GPT-5.6 Sol. ➤ Achiev
Announcing Artificial Analysis Intelligence Index v4.2 (opens the original)
Read excerpt
We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming + AA-Briefcase, our agentic knowledge work evaluation with a private test set + Surge’s GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prev
Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency (opens the original)
Read excerpt
Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic. Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3 ➤ Strong agentic performan
Kimi K3: second only to Fable 5 on AA-Briefcase (opens the original)
Read excerpt
Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task Last week Kimi (Moonshot AI) released Kimi K3, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, comparable to models such as Opus 4.8 and GPT-5.5. On AA-Briefcase, Kimi K3 scores an Elo of 1543, a +727 improvement over Kimi K2.6 and the second highest score recorded, behind only Claude Fable 5 (157
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
Top 100K sites
For Artificialanalysis, the outlet · Measured Aug 1, 2026
Website popularity band, not a count of readers or article views.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.