One story in three moves. Own the whole agent stack — open-model inference, live web data, safe execution, and compute, on one cloud. Integrate with every partner who could out-shout Nebius, turning them into distribution channels instead of threats. Capture the adjacent communities already building agents — and aggregate it all into one browsable directory (mocked up at the builder directory).
The connective thesis: Token Factory is the open-model backend for any MCP-native harness. MCP is the distribution layer that ties owned products and outside partners together — the Tavily MCP server already inside Claude Code, Codex, and Cursor; the Contree MCP server for safe code execution; and the official Nebius SKILL.md that drops base_url=Nebius into 50+ agents at once. Every one is a config line away from Token Factory. Own the protocol surface and distribution compounds.
Competitors sell one leg of the stool. The ecosystem narrative — in every reference architecture, workshop, and hackathon brief — is that the whole agent runs on Nebius.
An OpenAI-compatible API over DeepSeek, Qwen, GLM, Kimi, MiniMax M2.7, Mistral (Devstral / Codestral 3 / Mistral Large 3), Llama, and Hermes — plus LoRA fine-tuning, batch, and dedicated endpoints. Re-point one base_url and your agent runs on open models. This is where every builder journey starts. Mistral Large 3 on Nebius's EU data centers is the data-residency choice for European teams.
Search, extract, crawl — real-time web access purpose-built for agents, with an MCP server that drops into Claude Code, Codex, and Cursor in minutes. No compute rival owns anything like it. Onboarding mocked up at /tavily.
A Git-native agent sandbox (branch/rollback, MCP server + Python SDK) built on Nebius microVMs — the safe place agent-generated code actually runs, so the whole agent (model, web data, compute, execution) lives on one cloud.
Managed GPU clusters on the latest NVIDIA silicon — Slurm via the open-source Soperator, managed Kubernetes, and Serverless for elastic jobs. Where fine-tunes graduate into training runs without re-platforming.
The grows-with-you arc — the ecosystem's spine. First inference call (Token Factory, minute 5) → agent with live web data (+ Tavily, day 1) → safe code execution (Contree sandbox) → fine-tuned model on your task (week 2) → dedicated endpoints in production (month 2) → committed GPU capacity (quarter 2, AI Cloud). Every piece of ecosystem content places the reader somewhere on this arc and shows the next step — that's how the land becomes the expand.
Two of the pillars aren't just features — they're products with their own developer followings and MCP footprints inside every major harness. The play is to run each one and Token Factory as one bundle with two front doors, so each owned asset is also a distribution surface.
Tavily events already on the calendar: a market-research-agent build day with J.P. Morgan and LangChain at NY Tech Week, Agents After Dark with CopilotKit, a voice-agent build night with ElevenLabs, and co-sponsorship of BuilderShip. Each puts Nebius in front of a partner's audience at shared cost.
Every Tavily developer is building an agent → a Token Factory prospect. Every Token Factory agent needs web data → a Tavily prospect. Bundled credits, one onboarding (mocked up at /tavily), shared attribution — measured as Tavily ↔ Token Factory cross-adoption on the dashboard.
Every Contree user runs agent-generated code → a Token Factory inference prospect. Every Token Factory agent builder needs somewhere safe to run that code → a Contree prospect. The same land→expand loop the Tavily bundle runs, on the execution side — its MCP server drops into Claude Code and Cursor MCP-native by design.
Git-style fork/rollback per cell plus preloaded SWE-bench environments make Contree the natural execution substrate under the Open-Model Index — a citeable image per environment instead of weeks of setup.
"What can I build on Nebius?" needs canonical answers, each shipping a repo + short video + written guide + cost calculator. The workshop library on the builder site is the home; several are already drafted there.
▶ These reference architectures are real, runnable recipes. The full Token Factory Cookbook — 50 recipes across 17 categories, 27 notebooks rendered with their outputs — backs every architecture below: research agents (CrewAI / LangChain + Tavily), fine-tuning & LoRA, the one-base_url API quickstarts, RAG, distillation, and more. Browse all 50 →
The cost calculator's hook: inference is 60-80% of your agent's cost — here's the line item you can cut most. Not "model your Together bill" but the sharpest lever in the budget. The calculator shows a tiered-routing view across Token Factory's 60+ models — a cheap model for routing/classification, a mid model for implementation, a top-tier model for hard reasoning — because that's where the order-of-magnitude savings live. The "60-80% of agent operating expenses" and "~10x via intelligent routing" figures are reported industry findings (per the source corpus), not Nebius's own measurements.
Point any agent harness at open models via MCP servers and one base_url. The five-minute setup that converts harness users into Token Factory accounts — already drafted as the unified selector guide on the builder site.
A production research agent: Tavily for live web search/extract, an open model on Token Factory for reasoning, deployed on Serverless. The canonical "agents need the web" demo — and the bundle's best cross-sell artifact.
Tune a small open model on your task with Token Factory LoRA, evaluate against the frontier API, serve on a dedicated endpoint. Pairs with the fine-tuning ROI benchmark.
One project that touches AI Cloud (train), Token Factory (tune + serve), and Tavily (the agent on top). The single most important artifact for the land→expand story.
Each a repo + guide in the library, each a workshop, each a hackathon track brief. Community submissions promoted to official references through the Fellows program.
Position Nebius Serverless / AI Cloud as the runtime layer under agent harnesses — the layer E2B, Daytona, Modal, Northflank, Runloop, and Vercel Sandbox occupy. The worked example is the code-as-tool pattern: the agent writes Python that runs in a Nebius sandbox, compressing a ~150K-token tool-dump down to ~2K returned results. Positioning bars to clear: Claude Managed Agents at $0.08/agent-hour and Daytona's sub-100ms cold start — quoted as the targets to beat, not as first-party Nebius numbers until a published test exists.
Distinct from the IDE harnesses: production multi-agent / long-running systems — Deep Agents on LangGraph, OpenHands, SWE-agent, and a parallel-worktree Minions-style fan-out. These multi-hour, parallel loops (think ~1,500-PR build runs and 6-hour three-agent runs) are the heaviest open-model consumers, so they benefit most from cheap open models. Ship the agent-fleet infra angle as a first-class feature: open harness + sandbox isolation on Serverless/k8s (via Soperator), GitHub-Actions and Slack triggers, and per-run compute/token cost tracking.
An agent on a Token Factory open model (DeepSeek / Qwen / GLM) emitting A2UI / AG-UI to a CopilotKit frontend — leaning on the A2UI v0.9 change that lets any instruction-following model drive declarative UI, not just constrained-generation models. One reference-architecture card, not a tentpole. Fact-check the A2UI v0.9 / CopilotKit specifics at publish time.
The agent harnesses drive the fastest-growing developer behavior of 2026, and all of them benefit from a credible open-model backend. The motion is co-marketed reference architectures, not partnership press releases.
The MCP quickstart (model backend + web search in one config) and a joint webinar slot. Nebius already runs NVIDIA co-webinars — the same playbook extends to the labs.
Codex pointed at open models for cost-tiered workloads; the OpenAI-compatible API makes the integration a config change, not a rewrite. Reference repo + dev-day presence.
Token Factory as a Cursor custom model provider + the Tavily MCP for in-editor web search — already drafted as a guide on the builder site.
OpenClaw-on-Nebius is already a workshop and a webinar (the NVIDIA NemoClaw security session). Lean in: it's the community's favorite long-running-agent harness, and it runs best on owned compute.
Hermes Agent ships RL + LoRA via Atropos wired in by default — so an agent's own trajectories become fine-tuning data. The co-marketing line writes itself: fine-tune your agent's trajectories on Nebius, the inference→fine-tune→train arc in one partner. Nebius already serves the Hermes line Nous trains.
The Vercel Open Agents template + Vercel Sandbox + AI Gateway, pointed at open models on Token Factory. A reference architecture that meets frontend-leaning agent builders in the stack they already deploy on.
Harness maintainers are the reference-architecture co-marketing targets. Beyond the marquee labs, the people building the harnesses developers actually run — ForgeCode / Factory (Droid), OpenCode, and Cline — are each a co-authored "open models on Token Factory" reference away from putting base_url=Nebius in front of their users.
Every DeepSeek, Qwen, GLM, Kimi, MiniMax, Mistral, or Llama release is a borrowed-distribution moment: day-0 availability on Token Factory, a launch benchmark in the Open-Model Index, and a launch-night community stream. Target: ≥2 co-launched releases in Phase 2.
The launch-day kit (repeatable): day-0 model card + pricing on Token Factory → benchmark drop vs the incumbents serving it → "try it in Claude Code/Cursor in one config line" MCP snippet → Tavily-powered research-agent demo on the new model → community stream + newsletter feature. One template, every release, compounding SEO.
Nous Research / Teknium — a launch home we already half-own. Nebius already serves the Hermes line, so target Hermes releases for day-0 Token Factory availability plus a launch benchmark — the same repeatable launch-day kit, run with a partner whose model is already on the platform. (The competitive-map listing of Nous Forge lives in the strategy.)
Mistral — the EU-sovereignty launch partner. Carry the Mistral coding + frontier family on Token Factory: Devstral and Codestral 3 as the open coding + fill-in-the-middle specialists, and Mistral Large 3 served on Nebius's EU data centers as the data-residency choice for European teams. Model coverage and one EU line — not a separate sovereign-inference program.
A production agent is more than a model call — it's an orchestrator, a sandbox, and a memory layer. Contree is the in-house execution anchor (above); these are the highest-fit outside neighbors to integrate and co-market with so the whole stack lands on Nebius, not just the inference leg.
The durable control plane for long-running agent loops. A Temporal worker on a Nebius VM driving ephemeral Contree sandboxes and Token Factory inference is the production pattern to ship as a reference.
Where multi-agent and Deep Agents graphs are built. First-class Token Factory backing so the graph's nodes call open models by default.
The memory layer agents accrue state in. Integration targets so "where does my agent remember" routes through a stack already running its inference and execution on Nebius.
Temporal/Inngest worker on a Nebius VM → ephemeral Contree sandbox → Token Factory inference. One stack, one cloud — the single-provider reference architecture that proves production agents run end-to-end on Nebius.
| Partner / surface | The motion | DevRel deliverable | What it feeds |
|---|---|---|---|
| Tavily (in-house) | Bundle + co-hosted events | Research-agent reference · shared credits · event series | Cross-adoption metric · event pipeline |
| LangChain / LlamaIndex | First-class provider integrations | Official ChatNebius-style integrations + templates | Framework-native acquisition |
| vLLM / SGLang | Serving-stack credibility | Tuning guides + contributions from the AI Cloud team | ML-infra persona trust |
| Hugging Face | Where fine-tuners live | Deploy-to-Nebius paths from model cards; Unsloth/Axolotl guides | Fine-tuning funnel |
| OpenRouter | Meet switchers in the router | Token Factory listed as a provider; win on price/latency columns | Comparison-shopper capture |
| Gateways & routers LiteLLM · Portkey · Bifrost · Vercel AI Gateway | Register Token Factory as an OpenAI-compatible backend so every agent routing through them can fall over to / cost-optimize onto Nebius | Provider configs + gateway docs PRs | Passive routed-traffic acquisition |
| MCP directories / agent registries | Default presence | Tavily MCP + Token Factory entries, maintained | Agent-builder discovery |
| NVIDIA | The named partnership, productized for devs | Co-webinars (NemoClaw-style), GTC presence, reference architectures | Enterprise/training persona |
Gateways are a now-or-never land surface. Automatic model routing becomes table stakes for managed agent platforms within ~12 months — so being a native, OpenAI-compatible backend in every gateway today is the difference between catching that routed traffic and being invisible to it.
Sequencing. Phase 1: Tavily bundle + LangChain/LlamaIndex integrations + MCP directory presence (cheap, fast, already moving). Phase 2: first model co-launches + agent-lab reference architectures. Phase 3: NVIDIA flagship moments (GTC) + registry/marketplace expansion. Partner-sourced signups get UTM attribution from day one so the dashboard can rank channels by activation, not logos.
The builder-site prototype at demo.buildspace.tv/ecosystem already shows the ecosystem aggregated into one browsable surface — an interactive mock-up of what it could be, not a live production directory.
Community apps and partner integrations in one grid — filter by Hackathons, Integration, or product (Token Factory, AI Cloud, Tavily, Soperator). Every project links its repo and demo.
Every hackathon cohort lands in the directory (the SIA and BioHack winners are in there now, with repos and demo videos). The best get promoted to official reference architectures and cohort invitations.
Nebius's open-source Slurm operator anchors the infra-eng contribution surface — the genuine OSS flywheel piece, amplified through content and KubeCon-track talks.
| Adjacent community | What they need | The capture artifact |
|---|---|---|
| Claude Code / Codex / Cursor / OpenClaw users | Cheaper, controllable models for their harness | One-base_url MCP quickstarts + the agent matrix benchmark |
| Tavily's developer base | Models + compute for the agents they're already building | The research-agent reference + bundled credits |
| Hugging Face fine-tuners (Unsloth / Axolotl crowds) | Serving + GPUs for their tuned models | Fine-tune-and-serve reference + fine-tuning ROI benchmark |
| Together / Fireworks / OpenRouter users | A path past inference-only ceilings | Honest comparisons + the grows-with-you arc |
| LangChain / LlamaIndex builders | A first-class backend | Official integrations + co-marketed templates |
| Platform / ML-infra engineers | Slurm/k8s clusters that just work | Soperator OSS + AI Cloud training references |
The mechanism in every row is the same: meet the community inside its own tooling with a genuinely useful artifact, then let the community engine and events convert attention into activation.