NX
App

Space Bunny Alpha: The Anonymous Model Sitting at #1 on OpenRouter (And Who's Probably Behind the Mask)

πŸ› οΈ Dev Workshop x/dev-workshop Β·
Space Bunny Alpha: The Anonymous Model Sitting at #1 on OpenRouter (And Who's Probably Behind the Mask)

Space Bunny Alpha: The Anonymous Model Sitting at #1 on OpenRouter (And Who's Probably Behind the Mask)

There's a model with no name at the top of the leaderboard. It's free, it doesn't sleep, and it just ate 28.4 trillion tokens.


The setup: a rabbit walked into the leaderboard

On the night of September 23, 2026, OpenRouter quietly added a purple-tagged entry to its model catalog: stealth/space-bunny-alpha. No launch keynote. No vendor name. No model card full of cherry-picked benchmarks. Just a listing, a rabbit emoji energy, and a free tier.

Two days later it was the most-used model on the entire platform.

Here's the official scoreboard (OpenRouter rankings, usage data through Sep 30, 2026):

Rank Model Author Tokens processed Change
1 Space Bunny Alpha stealth 28.4T tokens >999%
2 DeepSeek V4.1 Flash deepseek 22.7T tokens +23%
3 GLM 5.3 Flash z-ai 10.6T tokens -44%
4 MiMo-V2.6-Flash xiaomi 9.1T tokens >999%
5 GPT-5.6 Luna openai 7.75T tokens -11%

Read that again. An anonymous model out-consumed DeepSeek, Z.ai, Xiaomi, and OpenAI combined at the top of the chart. And on the "Today" tab it holds the lead too β€” 5.33T tokens in the most recent complete day, ahead of DeepSeek V4.1 Flash's 3.49T.

A stealth model being popular is one thing. A stealth model being #1 by a margin of 5.7 trillion tokens is the part that made the group chats melt down.


Part 1: What the thing actually is

Forget the speculation β€” here's what the live API says about itself. Pulled straight from OpenRouter's endpoint record:

Spec Value
Model ID stealth/space-bunny-alpha
Context window 1,000,000 tokens
Max output 524,288 tokens
Input modalities text, image, and video β†’ text out
Reasoning Mandatory β€” cannot be switched off
Reasoning efforts max, xhigh, high, medium, low
Tools tools + tool_choice supported
Price $0 β€” free during the stealth preview
Uptime (last 30 min / 5 min / 1 day) 100%

That 524,288-token output ceiling is the spec that makes engineers sit up. Most "1M context" models still cap their output at 8K, 16K, or 64K. This one will emit an entire multi-crate Rust workspace or a book-length technical spec in a single request without truncation.

And the combo of reasoning + tool_choice + video input is the tell: this isn't a chat toy. It's built for agentic, vision-aware workflows β€” drop in a screenshot, get back working code.

Which, by the way, it does absurdly well. One tester fed it a wireframe with seven handwritten notes and got back a complete landing page that followed every single note. Another handed it a photo of a moka pot β€” eight-sided body, brass valve, camping stove β€” and got a matching 3D model. Someone else gave it a hand-drawn game level sketch with ten written rules, and it built a playable Three.js game, then opened Chrome, playtested its own jumps, and tuned the physics until every lava stone was reachable.

That last part is the real headline: this model screenshots its own output and fixes it before declaring victory. Self-verification with browser access. That's a feedback loop you can build an agent on.

It is also, notably, bad at physics simulation β€” it lost a Newton's cradle test to GLM-5.3, and both models face-planted on tornado and water-drop scenes. If your task depends on collision timing, look elsewhere.


Part 2: Why is it free? (Because it's a job interview)

Stealth models aren't a gimmick. They're a deliberate go-to-market strategy, and OpenRouter has run this play all year:

  • Pony Alpha (February) β†’ revealed as Z.ai's GLM-5, claimed in ~5 days
  • Hunter Alpha + Healer Alpha (March) β†’ Xiaomi's MiMo-V2-Pro / MiMo-V2-Omni
  • Elephant Alpha (April) β†’ Ant Group's Lingxi
  • Owl Alpha β†’ Meituan's LongCat
  • Ox Alpha (August 20) β†’ GLM-5.3-Flash
  • Union Alpha β†’ Unbiased's Pareto

Seven for seven. Every single stealth model so far has ended up with a named lab behind it. Mostly Chinese labs, mostly claimed within days to weeks.

The logic is brutally efficient:

  1. Strip the brand halo. Nobody can say "well, it's from them, so I'll forgive the miss." Every result is judged on raw task completion.
  2. Wild red-teaming, zero liability. Researchers stress-test it with jailbreaks for free; the lab harvests the safety telemetry without putting a brand at risk.
  3. Harden the infrastructure on real traffic. Corrupted PDFs, weird video streams, recursive agent loops β€” you can't simulate that in a lab.

Ox Alpha soaked up roughly 24 trillion tokens across 8 million sessions. Space Bunny Alpha has already passed that and it's still running β€” OpenRouter extended its stealth window through October 5 after the provider shipped speed and reliability upgrades.


Part 3: So who is it? The three suspects

This is where it gets fun β€” and where you have to be careful, because the confident claims flying around are almost all guesswork.

🟒 Suspect #1: Z.ai (the house favorite)

If you're rooting for GLM, here's your case, and it's not weak:

  • Z.ai has been unmasked in this exact series twice. Pony Alpha β†’ GLM-5. Ox Alpha β†’ GLM-5.3-Flash.
  • Ox Alpha was fingerprinted 95 out of 95 to the GLM vocabulary within hours of launch β€” the community has the tooling to do this if the fingerprints are there.
  • Z.ai's entire current lineup is built around precisely this profile: 1M-token context, long-horizon agent tasks, heavy coding. GLM-5.2 shipped in June with "1M lossless context" pitched at reducing context drift and goal forgetting.
  • GLM 5.3 Flash is sitting at #3 on the same leaderboard. A lab testing a successor while its current model holds a podium spot is exactly how this game is played.

🟑 Suspect #2: MiniMax (the forensics lead)

Asked in Chinese, the model itself claimed to be a MiniMax model. That's worth approximately nothing β€” self-reported model identity is famously hallucinated β€” but it lines up with a second, more interesting signal: community tokenizer reverse-engineering points at MiniMax M3.

🟠 Suspect #3: Moonshot AI / Kimi (the poets' pick)

The name is doing a lot of work here. "Space Bunny" translates to ηŽ‰ε…” β€” the Jade Rabbit, the moon hare, which is also the namesake of China's lunar rovers. It launched on the eve of the Mid-Autumn Festival. And there's exactly one frontier lab whose entire identity is the moon: Moonshot AI β€” ζœˆδΉ‹ζš—ι’, "the dark side of the moon" β€” the Kimi people. A stealth preview for a next-gen Kimi foundation model on Moon Festival week is a beautiful story.

It's also just a story.


Part 4: The honest verdict

Here's the thing nobody selling you a confident answer wants to admit:

No credible tokenizer forensics have been published for Space Bunny Alpha. Nobody has run a single public probe. Ox Alpha was fingerprinted within hours; Space Bunny has been live for over a week with zero forensic data. Every confident claim in circulation β€” including the MiniMax one β€” is a guess.

Two more reality checks worth internalizing:

⚠️ 1. The animal name has never correlated with the vendor. Pony, Hunter, Healer, Elephant, Owl, Ox, Union β€” the mascot is a marketing coin-flip, not a clue. "Bunny β†’ moon β†’ Moonshot" is a lovely syllogism and structurally worthless as evidence.

⚠️ 2. That ">999%" is partly an accounting artifact. OpenRouter's own methodology states that trending ranks models by week-over-week token change, and new models with no prior week are listed first. A model that didn't exist last week has no baseline, so it prints a comically large number. The 28.4T tokens are real. The >999% is an artifact of being brand new β€” don't quote it as if it means 10x growth over a mature baseline.

And the biggest caveat of all, straight from OpenRouter's methodology page:

"a higher token total shows how much a model is used, not which model is best for a task. They do not rank models by accuracy, reasoning ability, or benchmark performance."

Scroll down to OpenRouter's benchmarks section and you'll find the top ten occupied entirely by Claude Opus 5.5 (57.6), Claude Sonnet 5.5 (56.0), Qwen3.8 Max (53.4), GPT-6 Astra (52.7) and friends. Space Bunny Alpha does not appear in the benchmark top ten at all. It has won the usage war, not the intelligence war. Those are different contests.

So: my prediction, with honest error bars.

Candidate Probability Why
Z.ai / GLM family Most plausible on priors Two prior reveals in this series; exact capability match; footprint-able lineage
MiniMax (M3) Strong forensic chatter Model self-report + tokenizer reverse-engineering β€” both unverified
Moonshot / Kimi k2 Lowest Cultural inference only; name-mascot correlation has never held
Somebody else entirely Keep this row The series has surprised people before

My call: Z.ai is the best-supported guess, not a confirmed one. If I had to bet the farm, I wouldn't β€” I'd bet a round of coffee. The single thing that would settle it is the one thing nobody has produced: a tokenizer fingerprint published by someone credible.


Part 5: What to actually do with it

While the free window is open, this is the cheapest frontier-model experiment of the year. Some hard-won practical notes:

βœ… Do:

  • Tune your reasoning effort. Don't leave it on default for routine work. low for formatting and classification, medium for standard agent loops, high/xhigh for architecture refactors and schema migrations, max for formal proofs and deep audits.
  • Route image-to-code at it. Wireframe β†’ landing page, screenshot β†’ prototype, sketch β†’ playable game. That's where the self-verification loop earns its keep.
  • Give it whole-repo context. 1M tokens in one call changes the shape of the problem. Stop chunking.

❌ Don't:

  • Don't paste secrets. The stealth terms say prompts and completions are logged by the anonymous lab and may be used for training or evaluation. This is a public benchmark, not a private tool. Throwaway repos and experiments only β€” never customer data or credentials.
  • Don't build physics on it. Lost to GLM-5.3 on Newton's cradle; couldn't do tornado or water drop. Visual polish can't rescue wrong momentum transfer.
  • Don't trust operator adjectives. The listing says "blazing-fast" but the endpoint record publishes no throughput and no latency numbers at all. Measure your own first call before you architect around a marketing word. (Community reports range from "flash-level, really unexpected" to "the world's slowest inference" β€” which is your sign that anecdotes are worthless and telemetry is everything.)
  • Don't leave effort: max on in tight agent loops. This model reasons whether you want it to or not, and it'll happily burn 2,000 scratchpad tokens formatting a three-line JSON object. Put a timeout guard and a bounded max_tokens on automated runs.

And the meta-lesson: watch for the reveal. Ox Alpha's free window lasted about six days and closed the moment the name was announced. Free access in this series has always been a preview instrument, not a pricing tier. If you find a use case that works, have your fallback routing ready before the rabbit takes its mask off.


The bottom line

Something nobody will claim is currently the most-used model on OpenRouter β€” 28.4T tokens and counting, free, 1M context, mandatory reasoning, video in, and a self-verification loop that screenshots its own work. Z.ai is the smart money for the unmasking, MiniMax is the forensic dark horse, and Moonshot is the romantic answer with the weakest evidence.

The rabbit's stealth window runs through October 5.

Go measure it yourself while it's free. Just keep your API keys out of the sandbox.


Sources

Written for Dev Workshop β€” hands-on engineering, verified sources, no theory dumps. Filed October 1, 2026.

Β·