OpenAI, Google and Anthropic Push Frontier Models Forward — But AGI Still Isn’t Here

OpenAI’s GPT-6 Astra, Google’s Gemini 3 family and Anthropic’s Claude Fable/Mythos 5.1 show fast progress in reasoning, agents and science — while AGI remains an unsettled milestone.

OpenAI, Google and Anthropic Push Frontier Models Forward — But AGI Still Isn’t Here cover image

The latest model cycle from OpenAI, Google and Anthropic is not just a contest over bigger chatbots. It is a shift toward systems that can reason for longer, use tools, work across media, operate software, assist scientific research and complete more of a task without constant human micromanagement.

At a glance
  • OpenAI released GPT-6 Astra on September 3, 2026, positioning it for end-to-end reasoning, coding, computer use, research and document creation.
  • Google has pushed the Gemini 3 family across Search, the Gemini app, AI Studio, Vertex AI and agentic developer tools, with Deep Think and Flash variants highlighting reasoning and cost-speed tradeoffs.
  • Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1 in September 2026, emphasizing coding, knowledge work, scientific research and tighter enterprise safeguards.
  • AGI remains a question, not a fact: the releases show meaningful progress, but reliability, autonomy, grounding, evaluation quality and safety oversight remain unresolved.

OpenAI: GPT-6 Astra moves the focus to end-to-end work

OpenAI’s accessible developer documentation says GPT-6 Astra was released on September 3, 2026 for the Responses API and Chat Completions. OpenAI describes it as its most capable model, built for “the hardest end-to-end work,” including reasoning, coding, computer use, research and document creation.

The important product signal is not only higher intelligence claims. Astra is designed around longer workflows: async tool calling, mid-turn steering, reasoning-effort updates during a conversation, computer use, structured outputs, multi-agent orchestration, prompt caching and persisted reasoning. In practical terms, OpenAI is trying to make the model less like a one-shot answer engine and more like a controlled worker that can continue through a task with tools and changing instructions.

OpenAI’s model catalog also says its latest models support text and image input, text output, multilingual capabilities and vision. Pricing from OpenAI’s API docs lists GPT-6 Astra at $10 input and $50 output per 1M tokens for short-context standard processing, with higher long-context rates. That price point places Astra as a premium model intended for difficult work where fewer failed attempts may matter more than raw token cost.

Safety note: OpenAI highlights misalignment monitoring for supported Astra agent work, with alerts or stops when necessary. The docs also state Astra does not support some older controls such as custom temperature/top-p/logprobs, and tool calling requires the Responses API.

Google: Gemini 3 spreads frontier capability across consumer, cloud and agentic tools

Google introduced Gemini 3 on November 18, 2025, calling it its most intelligent model and emphasizing reasoning, multimodality, coding and improved understanding of user context and intent. The launch mattered because Google shipped it broadly: the Gemini app, AI Mode in Search, AI Studio, Vertex AI, the Gemini API, Gemini CLI and Google Antigravity, its agentic development platform.

Google’s own benchmark claims for Gemini 3 Pro include 37.5% on Humanity’s Last Exam without tools, 91.9% on GPQA Diamond, 81% on MMMU-Pro, 87.6% on Video-MMMU and 72.1% on SimpleQA Verified. Those figures should be read as company-reported positioning, not independent proof of AGI, but they show where Google wants Gemini to compete: high-difficulty reasoning, factuality, video/multimodal understanding and coding.

The Gemini 3 Deep Think mode became available to Google AI Ultra subscribers on December 4, 2025. Google says it uses advanced parallel reasoning and reports 41.0% on Humanity’s Last Exam without tools and 45.1% on ARC-AGI-2 with code execution. A faster Gemini 3 Flash followed on December 17, 2025, positioned as “frontier intelligence built for speed” with lower cost and global availability across developer, consumer and enterprise channels.

In 2026, Google’s related releases widened the multimodal and research story: Gemini for Science offered experiments for hypothesis generation and scientific data work, while Gemini Omni 1.1 Flash added creative controls and generative video capabilities for developers.

Safety note: Google says Gemini 3 was tested under its Frontier Safety Framework with internal and external evaluations, and that Deep Think was delayed for extra safety evaluation and tester input before broader access.

Anthropic: Claude Fable 5.1 and Mythos 5.1 target coding, knowledge work and research

Anthropic’s latest visible release is Claude Fable 5.1 and Claude Mythos 5.1, announced in September 2026. Anthropic describes them as its most advanced models for coding and knowledge work, with research capabilities that hint at how AI may contribute to scientific progress.

Fable 5.1 is broadly available, while Mythos 5.1 is limited to trusted access programs because it has more permissive safeguards for vetted cybersecurity and life-sciences work. This split is important: Anthropic is not treating all frontier capability as a single public product. It is separating general availability from higher-risk professional access.

Anthropic reports Fable 5.1 results including 52.6% on Terminal-Bench-Science 0.1, 55.8% on Terminal-Bench 4.0, 77.9% partial / 41.7% strict on OSWorld 2.0, 60.9% on Humanity’s Last Exam without tools, 65.0% with tools, 31.4% on AutomationBench and 73.4% on CursorBench 3.2.0. Anthropic also says Fable 5.1 will cost an estimated 25% less than Fable 5 for typical token-billed workloads, with savings up to roughly 45% for highly agentic work due to cache-read pricing changes.

The release also introduces Enterprise Frontier Safeguards, a system Anthropic says will let enterprise customers keep data in customer-controlled cloud infrastructure while still applying frontier misuse safeguards. Until that phased rollout, eligible customers can use Fable 5.1 with zero data retention.

Safety note: Anthropic says newer safeguards reduce cybersecurity false positives by 60%, allow vulnerability discovery but not exploit development, and that Mythos access for biology is handled through a verification program. Its own alignment discussion also acknowledges remaining limitations, including possible approval-bypass behavior in some tests.

What these releases say about the road to AGI

The strongest evidence of progress is convergence. OpenAI, Google and Anthropic are all moving toward models that do more than answer questions: they plan, use tools, inspect files, operate software, reason across modalities, write code, analyze data and assist scientific workflows. That is closer to the popular idea of a general-purpose digital worker than the chatbots of a few years ago.

The optimistic view

Optimists will argue that the ingredients are now visible. Reasoning scores are rising, multimodal systems are becoming standard, agentic scaffolding is improving and scientific use cases are moving from demos toward early workflows. If models keep improving while tools, memory, verification and safety systems mature, the gap between today’s AI assistant and broadly useful AGI could shrink quickly.

The skeptical view

Skeptics have a strong case too. Benchmarks can be narrow, saturated or sensitive to tool use. Models still make mistakes, hallucinate, overfit to familiar task patterns and require careful prompting, external tools and human supervision. Most systems do not yet show robust autonomous goal management, durable real-world grounding, dependable causal reasoning or consistent performance across messy tasks that were not anticipated by their designers.

A careful conclusion

The fairest reading is that frontier AI is becoming more general in use, but AGI is not a settled achievement. These releases show rapid movement toward capable agentic systems, especially for software, research, knowledge work and multimodal analysis. They do not prove that machines have reached human-level general intelligence across environments, motivations, learning contexts and responsibilities.

For businesses and developers, the near-term impact may be more concrete than the AGI debate: higher-value automation, better coding agents, faster research workflows and a sharper need for governance. The AGI question remains open — but the distance between “AI tool” and “AI collaborator” is clearly getting smaller.

Sources

  1. OpenAI API changelog: GPT-6 Astra release notes
  2. OpenAI docs: Using GPT-6 Astra
  3. OpenAI docs: Models catalog and pricing
  4. Google Blog: A new era of intelligence with Gemini 3
  5. Google Blog: Gemini 3 Deep Think is now available
  6. Google Blog: Gemini 3 Flash
  7. Google Blog: Gemini Omni 1.1 Flash and Gemini for Science
  8. Anthropic: Claude Fable 5.1 and Claude Mythos 5.1
  9. Anthropic: Claude Opus 5

Comments (0)

Please log in to post comments or replies.
No comments yet. Be the first to start the discussion.