GPT Image 2.5 Sunburst and Flare are now live on CometAPI โ†’
new/CometAPI research

Gemini 4 Pro Is Coming: Everything We Know From test, leaks & RSI Rumors

Gemini 4 Pro checkpoint with a 256k output token limit (vs. prior ~64k), high-effort compute, and possible 1.5M+ context ambitions.

CometAPI
AnnaAI model and API research team
Updated Sep 17, 2026 10 min read
Gemini 4 Pro Is Coming: Everything We Know From test, leaks & RSI Rumors
Use this pattern

Make the first API call.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_COMETAPI_KEY",
    base_url="https://api.cometapi.com/v1",
)

response = client.chat.completions.create(
    model="gpt-5-mini",
    messages=[{"role": "user", "content": "Build this workflow."}],
)

print(response.choices[0].message.content)

TLDR: Google has confirmed Gemini 4 as its โ€œmost ambitious pre-training run yet.โ€ Community leaks in mid-September 2026 point to an early internal checkpoint of Gemini 4 Pro (codenamed โ€œargonโ€) featuring a reported 256k output token limit, possible multi-million-token context, and high-effort reasoning modes.Unofficial timelines cluster around October 2026 after pre-training delays. RSI (recursive self-improvement) links remain speculative. Until official release, developers can access current frontier modelsโ€”including GPT-6 Astra, Gemini 3.8 Flash, and Claude Fable 5.1โ€”through unified, discounted APIs such as CometAPI.

Key Takeaways

  • Gemini 4 is real and in training. Google said in July 2026 that it had started its "most ambitious pre-training run yet" for Gemini 4.
  • Gemini 4 Pro has not been officially announced. The name, model ID, pricing, API availability, context window, and final specifications remain unconfirmed. Leaked screenshots and community reports describe a Gemini 4 Pro checkpoint with a 256k output token limit (vs. prior ~64k), high-effort compute, and possible 1.5M+ context ambitions.
  • Gemini 4.0 appears to have quietly surfaced on Arena.ai (apparently, Gemini 3.8 Flash is being routed to the new Gemini Pro within Arena)
  • Most credible unofficial targets now point to an October 2026 public rollout after pre-training refinements.
  • Claims of strong RSI for Gemini 4 remain unverified; related research (Dream-RSI) improves exploration strategies without weight updates.
  • Until Gemini 4 Pro is officially released, developers can already experiment with current frontier models through platforms such as CometAPI, including Gemini 3.8 Flash, GPT-6 Astra, and Claude Fable 5.1.

What Is Gemini 4 Pro?

Gemini 4 Pro is the rumored flagship โ€œProโ€-tier model in Google DeepMindโ€™s next major Gemini generation. It sits above the lighter Flash variants and is positioned for high-complexity work: advanced coding, multi-step reasoning, long-horizon agentic tasks, multimodal understanding, and enterprise-scale document or codebase analysis.

Unlike the iterative 3.x Flash releases that have continued through 2026, Gemini 4 is described by Google leadership as requiring substantially larger base models to remain competitive at the frontier. Sundar Pichai stated on the Alphabet Q2 2026 earnings call: โ€œFor the next generation of frontier, youโ€™re going to need much larger base models. We are now training Gemini 4, and weโ€™re being very ambitious with it.โ€

The โ€œProโ€ designation historically signals the highest-capability publicly available tier optimized for quality over pure speed or costโ€”similar to earlier Gemini Pro and 2.5 Pro models that led coding and reasoning leaderboards for extended periods.

What We Know About Gemini 4 Pro(As of September 17)

Confirmed Official Signals

Google publicly acknowledged Gemini 4 in July 2026. In the launch post for Gemini 3.6 Flash (and related Flash variants), the team stated they had โ€œalready started our most ambitious pre-training run yet, for Gemini 4.โ€ Alphabet CEO Sundar Pichai reinforced this on the Q2 earnings call, describing Gemini 4 as requiring โ€œmuch larger base modelsโ€ to stay at the frontier, with explicit emphasis on coding and autonomous agents. He expressed internal excitement about progress and confidence that the results would please external users when released.

No official model card, parameter count, context window, benchmarks, pricing, or API endpoint for Gemini 4 or Gemini 4 Pro has been published. The Flash series (including the September 2026 Gemini 3.8 Flash and 3.8 Live variants) continues to receive frequent updates while the larger Pro-tier work proceeds.

Leaked Checkpoint and Specs

Community reporting indicates Google has produced an early internal checkpoint of the upcoming Pro model. one zero-shot peacock svg demo is making the rounds. previous checkpoint was Argon 160 / Gemini 3.8 Flash on arena.Reported details include:

  • A 256k-token output limit (a large increase over the ~64k of prior Gemini models).
  • Generation under a โ€œHighโ€ thinking-effort setting that took approximately 2.4 minutes.
  • Association with the Gemini 4 Pro tier rather than another Flash iteration.

Capital Expenditure and the RSI Connection

In early August 2026, Google DeepMind executive Jasjeet Sekhon (chief strategy officer) publicly described the industryโ€™s massive capital expenditures on AI as a bet on recursive self-improvement (RSI)โ€”systems that can improve themselves or the processes that produce the next generation of models. He noted that current revenues do not yet justify the scale of spending and that RSI is becoming a central part of the investment thesis, while acknowledging it has not yet been achieved.

Separately, Google researchers published the Dream-RSI paper in mid-September 2026, demonstrating recursive improvement of an agentโ€™s exploration strategies (without updating model weights) using Gemini 3.x models. This is related research but not evidence that Gemini 4 itself incorporates strong weight-space RSI.
The RSI framing underscores the strategic scale of the Gemini 4 effort: Google is investing at a level consistent with a generational leap rather than an incremental upgrade.

Arena.ai / LM Arena Mentions and Routing Claims

Arena.ai (the crowdsourced human-preference platform formerly known as LMSYS Chatbot Arena / LMArena) is a frequent early-testing venue for Google models. New anonymous or labeled Gemini entries often appear there before official announcements, and the community treats them as signals of near-term releases.

Gemini 4.0 has โ€œquietly surfacedโ€ and that traffic labeled as Gemini 3.8 Flash is sometimes being routed to a new Gemini Pro checkpoint inside Arena battles.:

  • http://Arena.ai Suddenly Launched a New Variant "gemini-3.8-flash" This Morning on September 17
  • Official 3.8 Flash Was Released on September 2, Reappearance of the Same Name Is Extremely Unusual.

Gemini 4 Pro Is Coming: Everything We Know From test, leaks & RSI Rumors

What remains unconfirmed

As of September 17, 2026, there is no authoritative Google announcement confirming:

Gemini 4 Pro attributeCurrent status
Gemini 4 trainingConfirmed
Gemini 4 Pro nameUnconfirmed
First Pro checkpointLeak/rumor
October releaseLeak/estimate
Novemberโ€“December releaseExternal estimate
1.5M contextUnverified leak
10M contextUnverified claim
Specific benchmark scoresUnverified
GPT-6 Astra comparisonUnverified
Claude Fable 5.1 comparisonUnverified
RSI achievementUnverified
API model IDNot announced
Official pricingNot announced

The Relationship Between Gemini 4 Pro and RSI

Recursive self-improvement (RSI)โ€”systems that can meaningfully improve their own capabilities or the processes that produce the next generationโ€”has become a frequent topic of speculation around every major frontier training run.

  • The claim that "Google has achieved RSI, making its new model more powerful" originates from posts on X, but it has not yet been confirmed by testing or official reports.
  • Separately, Google researchers (with collaborators) published Dream-RSI in September 2026. The system allows an agent to recursively improve its own exploration and search strategies by โ€œdreamingโ€ alternative policies and retaining the best, without ever updating the underlying model weights. Experiments used Gemini 3.1 Pro and Gemini 3.7 Flash. This is meta-level improvement of the orchestration layer, not full weight-space RSI that would automatically produce a stronger Gemini 4 from Gemini 3.
  • Broader 2026 research (AIDEยฒ, various agent-loop papers, and economic analyses) shows bounded or component-level self-improvementโ€”better prompts, skills, memory retrieval, or evaluation harnessesโ€”but not full closed-loop weight-level RSI that produces successively stronger base models without human intervention.

In short: Google is clearly investing in self-improving agent systems, and the Gemini 4 training run is described as unusually ambitious. Strong claims that the model itself embodies solved RSI remain unproven and should be treated as rumor.

What Features are confirmed or expected for Gemini 4 Pro?

Building on the trajectory from Gemini 1.5 โ†’ 2.5 โ†’ 3.x and the leaked argon details, the following areas are the most plausible sites of meaningful progress. Each is framed as an H3-level expectation grounded in either leak data or reasonable extrapolation from prior generations.

Dramatically Expanded Output Capacity

The leaked 256k output token limit would allow generation of entire multi-file codebases, long research reports, or complex structured documents in a single pass. Prior Gemini models were more constrained on output length, often requiring iterative continuation. This change alone would simplify agentic coding pipelines.

Longer Context Windows and Better Retrieval

Reports of a possible 1.5Mโ€“2M token context (still undecided) would extend Googleโ€™s traditional strength in long-context handling. Combined with improved retrieval and reduced โ€œlost-in-the-middleโ€ effects, this would benefit legal, scientific, and large-codebase analysis.

Stronger Coding and Agentic Performance

Pichai explicitly called out coding and agentic coding as focus areas. Gemini 2.5 and 3.x already delivered large jumps on SWE-bench, LiveCodeBench, and WebDev Arena. Gemini 4 Pro is expected to push further on multi-file refactoring, autonomous bug fixing, tool-use reliability, and long-horizon planning.

Deeper Reasoning / Adaptive Compute

The โ€œHigh thinking effortโ€ mode that consumed 2.4 minutes on the argon checkpoint suggests continued investment in inference-time compute scaling (similar to Deep Think or extended thinking modes in the 3.x series). Expect more reliable multi-step logic, lower hallucination rates on hard problems, and configurable effort levels.

Multimodal and Native Tool Improvements

Gemini has long been natively multimodal. Further gains in video understanding, real-time visual reasoning, and seamless tool calling (background tools without breaking conversation flow) are consistent with the direction of 3.8 Live and related releases.

Efficiency and Serving Optimizations

Even with larger base models, Google typically pairs frontier Pro models with efficient Flash variants and speculative decoding. Expect continued progress on tokens-per-second and cost-per-quality.

What assistance can CometAPI provide while waiting?

Until Gemini 4 Pro ships, developers can already access current state-of-the-art models through CometAPIโ€™s unified, OpenAI-compatible endpoint. Notable options include:

  • GPT-6 Astra โ€“ OpenAIโ€™s flagship for complex reasoning, coding, computer use, and long-horizon agentic work (1,050,000-token context, up to 128k output). Available on CometAPI at a permanent discount versus list pricing.
    Model page: https://www.cometapi.com/models/openai/gpt-6-astra/
    Usage guide: https://www.cometapi.com/how-to-use-gpt-6-astra-api/
  • Latest Gemini 3.x Flash and Pro-class models (including 3.8 variants) for multimodal and agentic workloads.
  • Seamless switching between providers with a single API key and OpenAI SDK compatibility.
  • Gemini CLI integration documentation for local agentic coding workflows.

CometAPIโ€™s catalog also surfaces emerging models quickly, so Gemini 4 Pro (once released) can be expected to appear alongside the existing lineup under the same endpoint.

How Gemini 4 Pro Compares (Projected vs. Current Frontier)

CapabilityGemini 3.x / 2.5 Pro (Public)GPT-6 Astra (Public)Gemini 4 Pro (Leaked / Expected)
Context windowUp to ~1โ€“2M tokens1,050,000 tokensRumored 1.5Mโ€“2M
Max output tokens~64k128,000Leaked 256k
Reasoning / thinking modesDeep Think / Extendedlow โ†’ max effort levelsHigh-effort (2.4 min observed)
Coding / agentic focusStrong (SWE-bench gains)Flagship strengthExplicit priority + leak claims
MultimodalityNative (text/image/video/audio)Text + image inputExpected continuation + improvement
StatusProductionProduction (Sep 2026)Internal checkpoint, not released

Table notes: Gemini 4 Pro figures are drawn from the argon leak and community reports; they are not official Google specifications. GPT-6 Astra specs are from public documentation.

Conclusion

Gemini 4 Pro represents Googleโ€™s stated next step at the frontier: a larger-scale pre-training effort explicitly aimed at keeping pace withโ€”and ideally surpassingโ€”contemporaneous models from OpenAI, Anthropic, and others. The September 2026 argon checkpoint provides the first concrete (if still unofficial) glimpse of expanded output capacity and high-effort reasoning. Release timing appears to be tightening around October 2026, though post-training and evaluation remain variables.

RSI remains an active research direction rather than a confirmed feature of the model. In the meantime, the practical path for developers is clear: continue shipping with the strongest available systemsโ€”GPT-6 Astra, current Gemini 3.x models, and peersโ€”while monitoring official Google channels for the Gemini 4 announcement. Platforms such as CometAPI already aggregate these models under one key and one billing relationship, reducing friction both today and when Gemini 4 Pro eventually arrives.

Continue learning

Connect this article to the next decision.

View all topics
Published on Sep 17, 2026
Last updated Sep 17, 2026
5 views
Reviewed for clarity, source attribution and current API terminology.

Ready to cut AI development costs by 20%?

Start free in minutes. Free trial credits included. No credit card required.

Read More