Agents as a service

The cloud platform for AI agents.

Orca is a cloud platform for running AI agents in production. Your app starts a session with one call to the OpenAI Agents API; Orca runs the agent to the end in its own sandbox and streams the result back, with spend controls and usage on every run.

Start with $20 in free credits · Blog · For AI assistants

New accounts start with $20 in free credits. No card required.

Why Orca

Most of an agent feature is not the agent. Orca runs everything around it.

Each system is weeks of engineering, then maintenance forever. On Orca they are running before your first API call.

Your team ships features. We carry the pager.

How it works

Your product makes one API call. Orca handles everything between the request and the result.

  1. Build: Describe an agent, not a codebase. Define the agent once in a profile: model, prompt, skills, sandbox. Every run uses it. Author profiles in the dashboard, no code needed. Attach skills, capabilities, and MCP servers.
  2. Launch: One SDK call starts a run. Your backend, a script, or CI starts a run with one call. Orca handles queuing, retries, and lifecycle. Go, Python, and TypeScript SDKs (plus REST). Reusable agent profiles for any job.
  3. Execute: Agents work in a locked room. Each run gets its own files, permissions, and network rules. Agents never touch your internal systems. Per-agent file permissions. Network egress protection built in.
  4. Stream: Every event, live in your app. Tool calls, file writes, and progress stream back as they happen. Your users watch the agent work. Server-sent events, no polling. OpenTelemetry traces on every service.
  5. Measure: Every run priced to the cent. Every run is priced the moment it finishes. Spend caps stop runaway agents before the bill. Per-run cost attribution. Fail-closed monthly spend caps.

Features

Queues, sandboxes, spend controls, observability. One platform, not five side projects.

Orchestrate: Run many agents as one system

Runs, sessions, pools, and workflows as API calls. One integration for Claude, Codex, and Vercel AI.

Workspace: Real files, without the risk

Agents get a real filesystem. You keep control of every byte.

Govern: Spend, access, and egress under control

Who can run what. How much it spends. Where its traffic goes.

Observe: End-to-end observability, by default

Follow any run live, event by event. No black boxes.

Integrations

Your backend calls Orca with the OpenAI client. Live agent runs come back.

API reference: https://docs.orcapods.ai/reference/overview

Deploy your way

Start in minutes on Orca Cloud, or run Orca’s sandboxes on your own infrastructure. Your agent API stays the same.

We run it. You ship.

Launch on Orca Cloud with nothing to operate. Build, run and scale from one place.

Your infrastructure. Your terms.

Keep the work on your own hardware: Orca’s sandboxes on your servers or bare metal, on terms shaped around your team.

FAQ

Direct answers to the questions engineering teams ask first.

Is this only for coding agents?

No. An agent can be a support triager, a research analyst, an operations bot or a code reviewer. Orca runs them all the same way: a model, instructions, tools and skills.

Why not just call the model APIs directly?

You can, until you need sessions that survive a dropped connection, a sandbox per run, spend limits that hold and files you can download. Orca is the layer you would otherwise spend months building.

Do I need a new SDK?

No. Orca speaks the OpenAI Agents API. Point the openai package you already use at https://api.orcapods.ai/v1 and keep your code standard.

How does it fit into our existing app?

Like any backend dependency. Your app starts a session with one call, Orca runs the agent to the end, and the events stream back into your own UI.

What does it cost to run?

A plan includes machine time. On Orca credit, models cost what OpenRouter charges plus its 5.5% fee, with no Orca markup. On your own provider key, Orca charges nothing for tokens. Details on the pricing page.

Which models can I use?

Any model in OpenRouter’s catalog on Orca credit, or your own key for OpenAI, Anthropic, OpenRouter, Vercel or cheaperinference. The model name says who pays, and Orca never switches it for you.

How do you keep agents from doing damage?

Each session works in its own sandbox. Every paid call is checked against your balance first, and a cap on calls in flight stops a runaway agent.

How do we see what agents are doing?

Live. Every message, tool call and file streams back while the agent works, and a dropped stream picks up where it left off.

Can we run it on our own infrastructure?

Yes. We can run Orca’s sandboxes on your own servers or bare metal, with dedicated capacity, negotiated rates and an SLA. Talk to us and we reply within one business day.

Built by you, or your agents

Paste one prompt into Claude Code, Cursor, or Codex. It installs the CLI, signs in with a device code, and follows the guide to build your first agent.

Onboard me to Orca (orcapods.ai), the cloud platform that runs AI agents through the OpenAI Agents API.

1. Install the Orca CLI: `curl -fsSL https://orcapods.ai/install.sh | sh` (https://docs.orcapods.ai/cli/install).
2. Run `orca login`. It prints a one-time code and a URL: relay both so I can approve the sign-in on another device.
3. Add Orca as an MCP server for yourself, as https://docs.orcapods.ai/cli/mcp shows. In Claude Code: `claude mcp add orca -- orca mcp serve`.
4. Ask me what job I want done before you build anything. Then create the agent: a model, instructions, tools and skills. The model name says who pays: `orca/openrouter/<id>` runs on Orca credit, `<provider>/<id>` on my own key for that provider.

The docs are at https://docs.orcapods.ai; start with the quickstart.

Deploy now · Read the AI assistant guide