Initiatives 2 & 7 · Ecosystem & partners

The ecosystem Nebius owns, integrates, and captures.

One story in three moves. Own the whole agent stack — open-model inference, live web data, safe execution, and compute, on one cloud. Integrate with every partner who could out-shout Nebius, turning them into distribution channels instead of threats. Capture the adjacent communities already building agents — and aggregate it all into one browsable directory (mocked up at the builder directory).

Own: Token Factory · Tavily · Contree · AI Cloud Integrate: agent labs · model launches · gateways · NVIDIA Capture: 190+ projects in the directory prototype

The connective thesis: Token Factory is the open-model backend for any MCP-native harness. MCP is the distribution layer that ties owned products and outside partners together — the Tavily MCP server already inside Claude Code, Codex, and Cursor; the Contree MCP server for safe code execution; and the official Nebius SKILL.md that drops base_url=Nebius into 50+ agents at once. Every one is a config line away from Token Factory. Own the protocol surface and distribution compounds.

Own · the three pillars (plus execution)

Models. Web data. Compute. Execution. One vendor.

Competitors sell one leg of the stool. The ecosystem narrative — in every reference architecture, workshop, and hackathon brief — is that the whole agent runs on Nebius.

Token Factory

The open-model on-ramp

An OpenAI-compatible API over DeepSeek, Qwen, GLM, Kimi, MiniMax M2.7, Mistral (Devstral / Codestral 3 / Mistral Large 3), Llama, and Hermes — plus LoRA fine-tuning, batch, and dedicated endpoints. Re-point one base_url and your agent runs on open models. This is where every builder journey starts. Mistral Large 3 on Nebius's EU data centers is the data-residency choice for European teams.

Tavily

The web-data layer

Search, extract, crawl — real-time web access purpose-built for agents, with an MCP server that drops into Claude Code, Codex, and Cursor in minutes. No compute rival owns anything like it. Onboarding mocked up at /tavily.

Contree

The execution layer

A Git-native agent sandbox (branch/rollback, MCP server + Python SDK) built on Nebius microVMs — the safe place agent-generated code actually runs, so the whole agent (model, web data, compute, execution) lives on one cloud.

AI Cloud

The scale layer

Managed GPU clusters on the latest NVIDIA silicon — Slurm via the open-source Soperator, managed Kubernetes, and Serverless for elastic jobs. Where fine-tunes graduate into training runs without re-platforming.

The grows-with-you arc — the ecosystem's spine. First inference call (Token Factory, minute 5) → agent with live web data (+ Tavily, day 1) → safe code execution (Contree sandbox) → fine-tuned model on your task (week 2) → dedicated endpoints in production (month 2) → committed GPU capacity (quarter 2, AI Cloud). Every piece of ecosystem content places the reader somewhere on this arc and shows the next step — that's how the land becomes the expand.

Own · the two in-house partner engines

Tavily and Contree: products we already own, run as channels.

Two of the pillars aren't just features — they're products with their own developer followings and MCP footprints inside every major harness. The play is to run each one and Token Factory as one bundle with two front doors, so each owned asset is also a distribution surface.

Tavily · proof this quarter

The co-hosting machine

Tavily events already on the calendar: a market-research-agent build day with J.P. Morgan and LangChain at NY Tech Week, Agents After Dark with CopilotKit, a voice-agent build night with ElevenLabs, and co-sponsorship of BuilderShip. Each puts Nebius in front of a partner's audience at shared cost.

Tavily · cross-sell

Two funnels, one bundle

Every Tavily developer is building an agent → a Token Factory prospect. Every Token Factory agent needs web data → a Tavily prospect. Bundled credits, one onboarding (mocked up at /tavily), shared attribution — measured as Tavily ↔ Token Factory cross-adoption on the dashboard.

Contree · cross-sell

The execution-side loop

Every Contree user runs agent-generated code → a Token Factory inference prospect. Every Token Factory agent builder needs somewhere safe to run that code → a Contree prospect. The same land→expand loop the Tavily bundle runs, on the execution side — its MCP server drops into Claude Code and Cursor MCP-native by design.

Contree · reproducibility

SWE-bench as tags

Git-style fork/rollback per cell plus preloaded SWE-bench environments make Contree the natural execution substrate under the Open-Model Index — a citeable image per environment instead of weeks of setup.

Initiative 2 · compose

Reference architectures — canonical, maintained, co-marketed.

"What can I build on Nebius?" needs canonical answers, each shipping a repo + short video + written guide + cost calculator. The workshop library on the builder site is the home; several are already drafted there.

▶ These reference architectures are real, runnable recipes. The full Token Factory Cookbook50 recipes across 17 categories, 27 notebooks rendered with their outputs — backs every architecture below: research agents (CrewAI / LangChain + Tavily), fine-tuning & LoRA, the one-base_url API quickstarts, RAG, distillation, and more. Browse all 50 →

The cost calculator's hook: inference is 60-80% of your agent's cost — here's the line item you can cut most. Not "model your Together bill" but the sharpest lever in the budget. The calculator shows a tiered-routing view across Token Factory's 60+ models — a cheap model for routing/classification, a mid model for implementation, a top-tier model for hard reasoning — because that's where the order-of-magnitude savings live. The "60-80% of agent operating expenses" and "~10x via intelligent routing" figures are reported industry findings (per the source corpus), not Nebius's own measurements.

1

Token Factory + Claude Code / Codex / Cursor

The agent on-ramp · co-marketed with the agent labs

Point any agent harness at open models via MCP servers and one base_url. The five-minute setup that converts harness users into Token Factory accounts — already drafted as the unified selector guide on the builder site.

🍳 Runnable recipe: API quickstarts · Google ADK tool-calling
2

The research agent — Tavily + Token Factory

The flagship Tavily pairing

A production research agent: Tavily for live web search/extract, an open model on Token Factory for reasoning, deployed on Serverless. The canonical "agents need the web" demo — and the bundle's best cross-sell artifact.

3

Fine-tune and serve an open model

The answer to frontier-API bill fatigue

Tune a small open model on your task with Token Factory LoRA, evaluate against the frontier API, serve on a dedicated endpoint. Pairs with the fine-tuning ROI benchmark.

🍳 Runnable recipe: LoRA · Fine-tune Llama · Add a LoRA
4

Train it, tune it, serve it

The full-stack proof — all three surfaces in one walkthrough

One project that touches AI Cloud (train), Token Factory (tune + serve), and Tavily (the agent on top). The single most important artifact for the land→expand story.

🍳 Runnable recipe: Fine-tuning pipeline · Distillation
5

OpenClaw on Nebius · RAG on Nebius · serverless agents · batch pipelines

The long tail, community-extendable

Each a repo + guide in the library, each a workshop, each a hackathon track brief. Community submissions promoted to official references through the Fellows program.

6

Run your coding-agent's sandbox on Nebius

The agent-execution wedge · Serverless as the runtime layer

Position Nebius Serverless / AI Cloud as the runtime layer under agent harnesses — the layer E2B, Daytona, Modal, Northflank, Runloop, and Vercel Sandbox occupy. The worked example is the code-as-tool pattern: the agent writes Python that runs in a Nebius sandbox, compressing a ~150K-token tool-dump down to ~2K returned results. Positioning bars to clear: Claude Managed Agents at $0.08/agent-hour and Daytona's sub-100ms cold start — quoted as the targets to beat, not as first-party Nebius numbers until a published test exists.

🍳 Runnable recipe: OpenClaw on Nebius · Tool calling
7

Point your multi-agent harness at Token Factory open models

The high-token-volume "land" workload

Distinct from the IDE harnesses: production multi-agent / long-running systems — Deep Agents on LangGraph, OpenHands, SWE-agent, and a parallel-worktree Minions-style fan-out. These multi-hour, parallel loops (think ~1,500-PR build runs and 6-hour three-agent runs) are the heaviest open-model consumers, so they benefit most from cheap open models. Ship the agent-fleet infra angle as a first-class feature: open harness + sandbox isolation on Serverless/k8s (via Soperator), GitHub-Actions and Slack triggers, and per-run compute/token cost tracking.

8

CopilotKit + A2UI on Token Factory — give your agent a UI

Generative UI on an open model

An agent on a Token Factory open model (DeepSeek / Qwen / GLM) emitting A2UI / AG-UI to a CopilotKit frontend — leaning on the A2UI v0.9 change that lets any instruction-following model drive declarative UI, not just constrained-generation models. One reference-architecture card, not a tentpole. Fact-check the A2UI v0.9 / CopilotKit specifics at publish time.

Integrate · agent labs

Anthropic, OpenAI, Cursor, OpenClaw — meet the harnesses where they are.

The agent harnesses drive the fastest-growing developer behavior of 2026, and all of them benefit from a credible open-model backend. The motion is co-marketed reference architectures, not partnership press releases.

Anthropic

Claude Code + Token Factory + Tavily

The MCP quickstart (model backend + web search in one config) and a joint webinar slot. Nebius already runs NVIDIA co-webinars — the same playbook extends to the labs.

OpenAI

Codex + Token Factory

Codex pointed at open models for cost-tiered workloads; the OpenAI-compatible API makes the integration a config change, not a rewrite. Reference repo + dev-day presence.

Cursor

Custom model provider

Token Factory as a Cursor custom model provider + the Tavily MCP for in-editor web search — already drafted as a guide on the builder site.

OpenClaw

The community harness

OpenClaw-on-Nebius is already a workshop and a webinar (the NVIDIA NemoClaw security session). Lean in: it's the community's favorite long-running-agent harness, and it runs best on owned compute.

Nous Research

Hermes Agent — fine-tune the trajectories

Hermes Agent ships RL + LoRA via Atropos wired in by default — so an agent's own trajectories become fine-tuning data. The co-marketing line writes itself: fine-tune your agent's trajectories on Nebius, the inference→fine-tune→train arc in one partner. Nebius already serves the Hermes line Nous trains.

Vercel

Open Agents on Token Factory

The Vercel Open Agents template + Vercel Sandbox + AI Gateway, pointed at open models on Token Factory. A reference architecture that meets frontend-leaning agent builders in the stack they already deploy on.

Harness maintainers are the reference-architecture co-marketing targets. Beyond the marquee labs, the people building the harnesses developers actually run — ForgeCode / Factory (Droid), OpenCode, and Cline — are each a co-authored "open models on Token Factory" reference away from putting base_url=Nebius in front of their users.

Integrate · model labs

Be the launch-day home for open-model releases.

Every DeepSeek, Qwen, GLM, Kimi, MiniMax, Mistral, or Llama release is a borrowed-distribution moment: day-0 availability on Token Factory, a launch benchmark in the Open-Model Index, and a launch-night community stream. Target: ≥2 co-launched releases in Phase 2.

The launch-day kit (repeatable): day-0 model card + pricing on Token Factory → benchmark drop vs the incumbents serving it → "try it in Claude Code/Cursor in one config line" MCP snippet → Tavily-powered research-agent demo on the new model → community stream + newsletter feature. One template, every release, compounding SEO.

Nous Research / Teknium — a launch home we already half-own. Nebius already serves the Hermes line, so target Hermes releases for day-0 Token Factory availability plus a launch benchmark — the same repeatable launch-day kit, run with a partner whose model is already on the platform. (The competitive-map listing of Nous Forge lives in the strategy.)

Mistral — the EU-sovereignty launch partner. Carry the Mistral coding + frontier family on Token Factory: Devstral and Codestral 3 as the open coding + fill-in-the-middle specialists, and Mistral Large 3 served on Nebius's EU data centers as the data-residency choice for European teams. Model coverage and one EU line — not a separate sovereign-inference program.

Integrate · run the whole stack on Nebius

The orchestration and memory neighbors, as backend integrations.

A production agent is more than a model call — it's an orchestrator, a sandbox, and a memory layer. Contree is the in-house execution anchor (above); these are the highest-fit outside neighbors to integrate and co-market with so the whole stack lands on Nebius, not just the inference leg.

Orchestration

Temporal

The durable control plane for long-running agent loops. A Temporal worker on a Nebius VM driving ephemeral Contree sandboxes and Token Factory inference is the production pattern to ship as a reference.

Orchestration

LangGraph

Where multi-agent and Deep Agents graphs are built. First-class Token Factory backing so the graph's nodes call open models by default.

Agent memory

Letta & Mem0

The memory layer agents accrue state in. Integration targets so "where does my agent remember" routes through a stack already running its inference and execution on Nebius.

The pattern

Control plane → exec plane → inference

Temporal/Inngest worker on a Nebius VM → ephemeral Contree sandbox → Token Factory inference. One stack, one cloud — the single-provider reference architecture that proves production agents run end-to-end on Nebius.

Integrate · frameworks, registries & NVIDIA

First-class everywhere builders already are.

Partner / surfaceThe motionDevRel deliverableWhat it feeds
Tavily (in-house)Bundle + co-hosted eventsResearch-agent reference · shared credits · event seriesCross-adoption metric · event pipeline
LangChain / LlamaIndexFirst-class provider integrationsOfficial ChatNebius-style integrations + templatesFramework-native acquisition
vLLM / SGLangServing-stack credibilityTuning guides + contributions from the AI Cloud teamML-infra persona trust
Hugging FaceWhere fine-tuners liveDeploy-to-Nebius paths from model cards; Unsloth/Axolotl guidesFine-tuning funnel
OpenRouterMeet switchers in the routerToken Factory listed as a provider; win on price/latency columnsComparison-shopper capture
Gateways & routers
LiteLLM · Portkey · Bifrost · Vercel AI Gateway
Register Token Factory as an OpenAI-compatible backend so every agent routing through them can fall over to / cost-optimize onto NebiusProvider configs + gateway docs PRsPassive routed-traffic acquisition
MCP directories / agent registriesDefault presenceTavily MCP + Token Factory entries, maintainedAgent-builder discovery
NVIDIAThe named partnership, productized for devsCo-webinars (NemoClaw-style), GTC presence, reference architecturesEnterprise/training persona

Gateways are a now-or-never land surface. Automatic model routing becomes table stakes for managed agent platforms within ~12 months — so being a native, OpenAI-compatible backend in every gateway today is the difference between catching that routed traffic and being invisible to it.

Sequencing. Phase 1: Tavily bundle + LangChain/LlamaIndex integrations + MCP directory presence (cheap, fast, already moving). Phase 2: first model co-launches + agent-lab reference architectures. Phase 3: NVIDIA flagship moments (GTC) + registry/marketplace expansion. Partner-sourced signups get UTM attribution from day one so the dashboard can rank channels by activation, not logos.

Capture · the directory

The ecosystem directory — 190+ projects, mocked up today.

The builder-site prototype at demo.buildspace.tv/ecosystem already shows the ecosystem aggregated into one browsable surface — an interactive mock-up of what it could be, not a live production directory.

190+
Projects listed
community apps + hackathon cohorts
3
Hackathon cohorts imported
robotics · JetBrains · SIA + BioHack
20+
Partner integrations
Tavily, LangChain, vLLM, agent harnesses…
Mock-up
Submit-a-project flow
community pipeline, zero-backend
Surface

One directory, filterable

Community apps and partner integrations in one grid — filter by Hackathons, Integration, or product (Token Factory, AI Cloud, Tavily, Soperator). Every project links its repo and demo.

Flywheel

Hackathon → directory → reference

Every hackathon cohort lands in the directory (the SIA and BioHack winners are in there now, with repos and demo videos). The best get promoted to official reference architectures and cohort invitations.

OSS

Soperator + open tooling

Nebius's open-source Slurm operator anchors the infra-eng contribution surface — the genuine OSS flywheel piece, amplified through content and KubeCon-track talks.

Capture · the map

Whose ecosystem we capture, and with what.

Adjacent communityWhat they needThe capture artifact
Claude Code / Codex / Cursor / OpenClaw usersCheaper, controllable models for their harnessOne-base_url MCP quickstarts + the agent matrix benchmark
Tavily's developer baseModels + compute for the agents they're already buildingThe research-agent reference + bundled credits
Hugging Face fine-tuners (Unsloth / Axolotl crowds)Serving + GPUs for their tuned modelsFine-tune-and-serve reference + fine-tuning ROI benchmark
Together / Fireworks / OpenRouter usersA path past inference-only ceilingsHonest comparisons + the grows-with-you arc
LangChain / LlamaIndex buildersA first-class backendOfficial integrations + co-marketed templates
Platform / ML-infra engineersSlurm/k8s clusters that just workSoperator OSS + AI Cloud training references

The mechanism in every row is the same: meet the community inside its own tooling with a genuinely useful artifact, then let the community engine and events convert attention into activation.