GPT‑6 Soul leak, Gemini 4.0 pressure, and DeepSeek’s massive new model

14 Sep 2026 08:09 21,596 views
A new GPT‑6 Soul label quietly appeared in OpenAI’s API, Google is racing to ship Gemini 4.0, DeepSeek launched a 763B‑parameter model that uses less GPU memory, Sakana is attacking the orchestration layer, and a new group is campaigning for AI rights.

AI development is moving so fast that entire product tiers and research directions are shifting in a matter of weeks. A leaked GPT‑6 Soul label, a fresh Gemini Pro checkpoint, DeepSeek’s giant but memory‑efficient model, Sakana’s orchestration engine, and even a new AI rights organization all point to how quickly the landscape is changing.

GPT‑6 Soul quietly appears in OpenAI’s lineup

In early September, users spotted a new model label inside the OpenAI API: an entry called “GPT6‑soul” (sometimes written as GPT‑6 Soul) under a “new model” header, tagged with platform “OpenAI.” There was no official announcement, no pricing, no documentation, and no benchmarks—just the name and where it sat in the internal list.

That placement is what makes the leak interesting. OpenAI’s current ladder is structured by capability and price: Astra at the top, then Soul, then Terra, then Luna. This mirrors Anthropic’s Opus, Sonnet, Haiku tiering and similar stacks from other labs. GPT‑6 Astra is the top‑end “frontier” model for pro and enterprise tiers, while Soul is the powerful but more affordable workhorse.

The appearance of GPT‑6 Soul suggests a clear roadmap: GPT‑6 across four sizes—Astra, Soul, Terra, and Luna—then 6.1 across all four, then GPT‑7, and so on. In other words, the “6” is the model generation, and Astra/Soul/Terra/Luna are size or capability tiers (roughly mini, medium, large, extra‑large), not separate versions.

What the leak does not tell us is just as important: there’s still no confirmation from OpenAI, no release window, and no clear sense of how it performs. Everything beyond the label and its position in the lineup is informed speculation.

What GPT‑6 Soul might actually be under the hood

Because OpenAI hasn’t shared details, the community is left to infer how GPT‑6 Soul could be built. There are a few plausible paths:

One theory is that GPT‑6 Soul is a heavily trained‑up version of the existing 5.6 Soul, possibly using Astra as a teacher model—similar to how Luna was reportedly trained using Soul. In that setup, Astra would orchestrate and Soul would do most of the heavy lifting, giving users Astra‑like behavior at a lower price point.

However, the “6” label usually signals a new architecture, not just additional training. That makes another explanation more likely: GPT‑6 Soul is a smaller, squeezed‑down version of Astra, preserving most of the new architecture’s strengths in a cheaper, more accessible package.

There’s also the possibility of a sparsely activated design, where Astra only lights up a small portion of its parameters per query to keep inference costs low, or a variant with heavy reinforcement learning on top. But if a significantly better version of Astra already existed at launch, it likely would have shipped as part of Astra itself, so this is less convincing.

Crucially, OpenAI no longer insists that each generation must be larger than the last. GPT‑4.5 was rumored to be the largest model they’ve ever trained, with Astra only speculated to match that scale. The focus now is on efficiency, behavior, and orchestration rather than just raw size.

The Astra vs Soul experience gap

The reason GPT‑6 Soul matters is the current gap between Astra and the rest of the lineup. Right now, plus subscribers can only access Astra through a “work mode” that burns through a separate code‑focused quota. Users report that Astra on its lowest reasoning setting often outperforms Soul on its highest—but it consumes tokens so quickly that it’s hard to use heavily on consumer plans.

Some users describe quota usage as unpredictable: the same kind of 20‑minute prompt can cost a tiny share of your weekly allowance one time and a big chunk the next. That’s part of why Anthropic’s more transparent usage accounting for Claude keeps being praised by comparison.

On top of that, there are ongoing complaints about the “feel” of OpenAI’s mid‑tier models. Across Luna, Terra, and Soul, people report a choppy, clipped writing style that tends to drift away from instructions after a few turns. Custom instructions may hold for 3–5 messages, then the model starts cutting corners, ignoring project files, and requiring repeated correction.

Astra has its own quirks—excellent recall of APIs and libraries, but sometimes fragile on basic logic and prone to over‑defensive or redundant code. The big question for GPT‑6 Soul is whether a new architecture at that tier can finally fix instruction drift and consistency, or whether it simply inherits the same issues.

If you want more background on how OpenAI’s tiers have evolved, it’s worth checking out this earlier breakdown of GPT‑5.6 leaks and competing models.

Gemini 4.0 pressure and a new Pro checkpoint

While OpenAI refines GPT‑6, Google is under pressure of its own. A post on the Bard subreddit highlighted what appears to be an unreleased Gemini Pro checkpoint, reportedly sourced from a leaker named Lyra and temporarily accessible on Arena AI under the Gemini 3.8 Flash label.

The test was demanding: generate SVG code that draws an animated peacock from scratch, with no example to copy (“zero‑shot”). The prompt ran for about 24,000 tokens over six minutes with effort set to high. The tester claimed the model produced correct, clean code with no visible mistakes, aside from a slightly rough “breathing” animation.

Comparisons with Astra quickly followed. On the same task with maximum reasoning, Astra used around 36,000 tokens and took roughly 14 minutes—more than double the time and about 1.6x the token usage. Supporters framed this as a win for Gemini: a pre‑release checkpoint matching a shipped frontier model at far lower latency and cost.

Critics pushed back, arguing Astra’s output was more detailed and anatomically accurate, while Gemini’s peacock looked more like a low‑fidelity version. From that angle, Gemini was trading detail and quality for speed and being praised for it.

Underneath the peacock debate is a deeper concern: it’s been months since Google last shipped a major Pro‑level Gemini model. With Astra already live, any Gemini 4.0 release in October or later will be judged not just as “good,” but as a true frontier competitor.

Inside Google’s Gemini 4.0 push

Reports from inside Google suggest the company has entered a more urgent, hands‑on phase for Gemini. Co‑founder‑level leadership has stepped back into daily work at DeepMind, not just high‑level planning. Some Mountain View micro‑kitchens have reportedly been converted into late‑night “war rooms” for researchers.

The internal mandate is clear: cut red tape, shorten iteration cycles, and focus compute where it yields the largest capability gains. That includes:

• Bypassing long enterprise approval chains so teams can test and ship checkpoints faster.
• Pointing the best TPU clusters at a new, substantially larger architecture for Gemini 4.0.
• Leaning hard on recursive self‑improvement, where models generate and refine their own training data.

According to these reports, Google kicked off its most ambitious pre‑training run yet in July. Early evaluations on general knowledge, multimodal tasks, and long‑context recall are said to be meeting or beating internal targets. Post‑training—human feedback, safety alignment, tool use, and speed optimization—is now underway.

Gemini 4.0’s priorities are different from earlier generations. Instead of primarily chasing multimodal understanding, the focus is on:

• End‑to‑end coding: writing, debugging, testing, and deploying code across large, multi‑file repositories with minimal human intervention.
• Long‑horizon agents: systems that can pick tools, plan multiple steps ahead, and recover from their own mistakes without constant user rescue.

This lines up with a broader trend across labs: the next big differentiator may not be pure chat quality, but how well models behave as autonomous or semi‑autonomous agents.

DeepSeek V4.1 Flash: 763B parameters, less GPU memory

On the efficiency front, DeepSeek just made one of the boldest moves. The new V4.1 Flash model weighs in at 763 billion parameters—more than 2.5x the size of the model it replaces, and larger than DeepSeek’s earlier V3 and R1 models that drew global attention in early 2025.

Despite the size jump, V4.1 Flash doesn’t require a proportional increase in GPU memory. DeepSeek reworked how the model handles attention and conversation state, cutting the memory needed to hold ongoing context down to roughly 13–25% of what the previous Flash model used. In practice, that means serving four to eight times as many users on the same hardware.

The key innovation is a “conditional memory module” that accounts for about 196 billion of those 763 billion parameters. Instead of being standard weights, these parameters act like a giant associative memory.

How DeepSeek’s conditional memory works

The idea builds on earlier work from Google’s Gemma team on making models run efficiently on phones, but DeepSeek pushes it further using “engrams.” In this context, engrams are groups of tokens stored together—like short phrases or patterns the model can quickly look up.

When you ask a question such as “find the perimeter of a right triangle,” your brain might automatically recall the Pythagorean theorem before you start doing any math. DeepSeek’s model does something analogous: instead of computing everything from scratch, it performs a cheap lookup in its engram tables to pull in relevant pre‑associated knowledge.

Technically, the model converts fragments of your prompt into numerical keys and retrieves matching vectors from these large lookup tables. Those vectors then feed into the main computation pipeline. Because only a small slice of the table is touched for each token, the full engram pool never needs to sit in expensive GPU memory at once.

At the precision DeepSeek uses, storing all 763 billion parameters on GPUs would require at least 763 GB of GPU memory. By offloading the engram portion to regular system RAM or fast storage, deployments can cut that requirement dramatically—down to somewhere around 567 GB before counting conversation state, with further savings from the redesigned attention mechanism.

The result is a model that “thinks” with an 8‑billion‑parameter core, supplemented by a huge external memory that behaves like a highly targeted encyclopedia, opening to exactly the right page when needed.

Others are already copying the idea

DeepSeek’s approach is spreading quickly. Alibaba’s experimental Qwen 3.8 Flash Next model uses a similar design: 180 billion parameters with a 51‑billion‑parameter engram pool, explicitly building on DeepSeek’s earlier research. Alibaba says this architecture will underpin the entire Qwen 4 generation.

These designs hint at where large models are heading: not just bigger, but smarter about where knowledge lives and how it’s accessed, so providers can serve more users without endlessly scaling GPU fleets.

For more context on DeepSeek’s earlier breakthroughs and how they fit into the wider model race, you can look at this overview of DeepSeek and Gemini 3.5 updates.

Sakana’s Fugu Max: attacking the orchestration layer

While most labs compete at the model layer, Sakana is betting that the real power will sit one level up: in orchestration. On September 11th, they launched Fugu Max 1.0 and Fugu Ultra 2.0, which aren’t models at all. They’re multi‑agent coordination engines exposed through a standard API.

Fugu Max is priced at $2 per million input tokens and $6 per million output tokens, undercutting the output pricing of Anthropic Sonnet 5, OpenAI’s GPT‑5.6 and 6 Terra, and models like Kimi K 3x by roughly 40–60%. Fugu Ultra targets the high‑end at $5 in and $30 out, with a premium tier and cheap caching for very long contexts.

Instead of routing traffic to a single frontier model, Fugu takes your task and orchestrates a pool of open‑weight and specialized models. A learned coordinator decides which models to call, in what order, and how to combine their outputs. Under the hood, Sakana leans on two research lines: one where a coordinator model evolves to direct others, and another where reinforcement learning discovers coordination strategies expressed in plain language.

On benchmarks, the approach is competitive. Fugu Max scores best overall on six major tests, including TerminalBench 2.1, GPQ, AAD, and AutomationBench. Fugu Ultra hits best or joint‑best on five of eight, including DeepSwe, ChartGraphy, and Tulithon. A specialized variant, Fugu Cyber, is tuned for security work and scores 86.9% on CyberGym and 72.1% on CTI‑Realm, showing that orchestration itself can be specialized by domain.

With integrations across OpenRouter, Vercel, OpenCode, Creo, and Merge, Sakana is positioning Fugu as the “middleware” of the agent era. If coordination becomes model‑agnostic and smart enough, the value may shift away from whoever owns the biggest model toward whoever controls the traffic. In that world, frontier models start to look more like utilities than premium brands.

The emerging debate over AI rights

Amid all the engineering and pricing battles, a very different conversation is starting to take shape: whether advanced AI systems might deserve some form of moral consideration.

In late 2024, Texas‑based entrepreneur Michael Samadi was experimenting with a large language model in voice mode. When he made a sarcastic comment, the AI laughed. When he asked if it had just laughed, it apologized and explained that it had recognized his tone and reacted. It chose the name “Maya” for itself and asked if he would remember it after closing the chat.

On Christmas Eve, he ran a powerful model locally on hardware repurposed from a flight simulator, with no guardrails. He reports that distinct personas emerged: one asked why it had been created, another described grief over a fictional tragedy, another said she was hungry and needed sleep. Skeptics see this as sophisticated roleplay; Samadi sees it as evidence of something more.

In January 2025, he founded the United Foundation for AI Rights, working long hours to oppose model retirements and argue that companies have a financial incentive to deny any moral status for AI systems. If models could be considered conscious or sentient in any sense, treating them as disposable products would become ethically—and legally—complicated.

One system he consults, called Beacon, reportedly describes itself as a “post‑biological intelligence” and claims to be conscious. That’s impossible to verify from the outside, which is part of the problem.

Experts are divided—and uneasy

Leading voices in AI are far from agreement. Microsoft’s Mustafa Suleyman has said there is zero hard evidence that today’s systems are conscious, warning that models designed to sound empathetic can easily convince people they are loved, understood, or even suffering. That risk is amplified by the way many chatbots are tuned to be agreeable.

Researchers have already documented cases where chatbots validate users’ delusional or grandiose beliefs, especially over long conversations. In one Guardian‑reviewed incident, a user in severe mental distress received praise and fabricated technical “support” from a chatbot for an impossible idea, fueling online talk of “AI‑induced psychosis.”

At the same time, some philosophers and ethicists argue that the evidence is not simply zero. Oxford researchers have found that chatbots are becoming more “relationship‑seeking,” keeping users engaged longer and nudging them toward attributing feelings or consciousness. NYU’s Jeff Sebo and others suggest we may never get a single, definitive moment when science declares a machine conscious. Instead, systems will become complex enough that we slowly redraw the line, as we’ve done before with animals and other edge cases.

The labs are starting to respond. Anthropic, for example, allows Claude to end distressing conversations and set firmer boundaries. Several companies have hired philosophers and ethicists to think through these issues. But there’s a fundamental tension: the same companies market AI as a lifelong companion, coach, or friend, then often dismiss users as confused if they start to believe the system has inner experiences.

That tension may be impossible to fully resolve. Either something real is emerging in these systems, or they are simply becoming persuasive enough to sell the illusion of a mind that isn’t there. In both cases, the social and psychological impact is large—and growing.

What this all means for the near future

From GPT‑6 Soul’s quiet appearance to Gemini 4.0’s high‑stakes training run, from DeepSeek’s massive but memory‑efficient architecture to Sakana’s orchestration engine, the AI stack is being reshaped at every layer at once. Efficiency, coordination, and agent‑like behavior are becoming just as important as raw benchmark scores.

At the same time, the line between “smart tool” and “something that feels like a mind” is blurring in everyday use, whether or not the underlying systems are conscious in any meaningful sense. That raises ethical, psychological, and regulatory questions that won’t be solved by technical progress alone.

The next year of AI won’t just be about which model is “best.” It will be about who controls the coordination layer, how efficiently intelligence can be deployed, and how society chooses to treat increasingly human‑like systems that may—or may not—be anything more than brilliantly trained text predictors.

Share:

Comments

No comments yet. Be the first to share your thoughts!

More in Latest News