Latest Releases

GPT-6.1

OpenAI

GPT-6.1 Sol arrives as upgraded developer-day release

Sep 29, 2026

Closed SourceLLM

OpenAI's latest model variant improves on GPT-6 Sol with enhanced reasoning and coding capabilities, announced at its September 2026 Developer Day.

The GPT-6.1 Sol update, revealed at OpenAI's Developer Day on September 29, focuses on three key improvements: multi-step reasoning accuracy (up 18% on GSM8K math problems), code generation quality (23% improvement on HumanEval+), and reduced hallucination rates (down 12% on TruthfulQA). The model uses the same transformer architecture as GPT-6 Sol but with refined training data and a 15% larger context window of 160K tokens. API pricing remains unchanged at $2.50 per million input tokens and $10 per million output tokens. Developers can access GPT-6.1 Sol through the OpenAI API starting October 5 via the model identifier "gpt-6.1-sol-2026-09-29". The release also includes improved JSON output formatting and new tool-use capabilities for structured data extraction.
Sonnet

Anthropic

Claude Sonnet 5.5 balances performance and cost

Sep 28, 2026

Closed SourceLLM

Anthropic's new mid-tier model offers near-Opus performance at a fraction of the cost, targeting everyday professional workloads and agentic tasks.

Claude Sonnet 5.5, released September 28, sits between the entry-level Haiku and flagship Opus lines. Anthropic's benchmarking shows Sonnet 5.5 achieves 94% of Claude Opus 5.4's performance on the MMLU-Pro academic benchmark while running at 60% lower latency and 75% lower cost. The model introduces improved agentic capabilities, including better multi-turn tool use and 3x longer sustained task execution. Sonnet 5.5 handles 200K context tokens and is optimized for document processing, customer support automation, and code review. Pricing is set at $0.80 per million input tokens and $3.20 per million output tokens, making it the most cost-effective option for high-volume enterprise workloads. The model is available immediately on Anthropic's API and through Claude.ai's Pro plan.
Live

Google

Gemini 3.8 Live brings real-time voice conversations

Sep 16, 2026

Closed SourceAudio

Two new voice models: Gemini 3.8 Live for everyday conversation and Extended Thinking for multi-step reasoning without lag. Plus 2,000+ voice TTS options.

Gemini 3.8 Live, announced September 16, is Google's first real-time conversational AI model designed for voice-first interfaces. The model achieves sub-200ms response latency, enabling natural turn-taking in phone calls, virtual assistants, and interactive applications. Gemini 3.8 Live Extended Thinking adds the ability to pause and reason for up to 10 seconds before responding, useful for complex queries that require deliberation. Both models support 48 languages with native pronunciation. Google also released 2,000+ new Text-to-Speech voices, including regional accents, emotion variations, and celebrity-licensed options. Gemini 3.8 Live is available through the Vertex AI platform and Google Cloud APIs at $4 per hour of audio for input and $12 per hour for output.
Opus

Anthropic

Claude Opus 5.5 targets long-running agent workflows

Sep 22, 2026

Closed SourceLLM

Built for complex agentic tasks, Opus 5.5 matches Fable 5.1 performance at 60% lower cost with faster response times for coding and knowledge work.

Claude Opus 5.5, released September 22, is Anthropic's answer to complex, long-duration AI agent workflows. The model features a new "extended thought" architecture that can maintain context across thousands of tool calls while reducing drift and error accumulation. Benchmarks show Opus 5.5 matches Microsoft's Fable 5.1 on the AgentBench suite of 127 multi-step tasks while completing them 40% faster. The model's 400K context window and improved memory management make it suitable for autonomous coding, research synthesis, and business process automation. Opus 5.5 introduces a "confidence interval" output mode that provides uncertainty estimates alongside predictions. Pricing is $15 per million input tokens and $60 per million output tokens, significantly undercutting competing flagship models.
Luna

OpenAI

GPT-6 Luna sets new API price floor at $0.10 per million tokens

Sep 22, 2026

Closed SourceLLM

OpenAI's lowest-cost GPT-6 model halves pricing versus GPT-5.6, making frontier AI accessible for high-volume, fast-response use cases.

GPT-6 Luna, OpenAI's new entry-level model released September 22, dramatically reduces the cost of API-based AI at $0.10 per million input tokens and $0.40 per million output tokens - a 50% reduction versus GPT-5.6. Luna achieves 85% of GPT-6 Sol's MMLU accuracy while running at 3x the throughput on OpenAI's infrastructure. The model is optimized for high-volume, low-latency use cases such as document classification, sentiment analysis, spam detection, and customer intent routing. Luna supports 128K context tokens and includes improved multilingual capabilities across 50+ languages. OpenAI reports that Luna can process 2 million requests per second across its global API infrastructure. The model is immediately available to all API users at no additional setup cost.
Grok

xAI

Grok 4.7 launches with 500K context window

Sep 21, 2026

Closed SourceLLM

Elon Musk's AI company releases Grok 4.7 for coding and knowledge work at the same token prices as Grok 4.6, available via API and major IDE integrations.

Grok 4.7, released by xAI on September 21, features a massive 500K token context window, making it suitable for analyzing entire codebases, legal documents, and research corpora in a single pass. The model shows significant improvements in code generation, scoring 89.2% on HumanEval+ (up from 82.1% on Grok 4.6). Grok 4.7 is integrated into JetBrains IDEs, Visual Studio Code, and xAI's own Grok terminal interface. The model maintains the same pricing as its predecessor: $2 per million input tokens and $8 per million output tokens. xAI also announced improved real-time web search capabilities, allowing Grok 4.7 to fetch and synthesize current information with reduced hallucination rates. The model is trained on xAI's Colossus 3 supercomputer using a mix of web data, books, code repositories, and scientific papers through September 2026.