Agents as a service
The cloud platform for AI agents.
Orca is a cloud platform for running AI agents in production. Your app starts a session with one call to the OpenAI Agents API; Orca runs the agent to the end in its own sandbox and streams the result back, with spend controls and usage on every run.
Start with $20 in free credits · Blog · For AI assistants
New accounts start with $20 in free credits. No card required.
Why Orca
Most of an agent feature is not the agent. Orca runs everything around it.
Each system is weeks of engineering, then maintenance forever. On Orca they are running before your first API call.
- Run queue, retries, and session state
- Sandbox isolation and provisioning
- File permissions and workspace storage
- Live event streaming into your app
- Usage metering and spend caps
- Rate limits, secrets, and egress control
- Audit logging and RBAC
- Traces, logs, and metrics
- Keeping up with new models and SDKs
Your team ships features. We carry the pager.
How it works
Your product makes one API call. Orca handles everything between the request and the result.
- Build: Describe an agent, not a codebase. Define the agent once in a profile: model, prompt, skills, sandbox. Every run uses it. Author profiles in the dashboard, no code needed. Attach skills, capabilities, and MCP servers.
- Launch: One SDK call starts a run. Your backend, a script, or CI starts a run with one call. Orca handles queuing, retries, and lifecycle. Go, Python, and TypeScript SDKs (plus REST). Reusable agent profiles for any job.
- Execute: Agents work in a locked room. Each run gets its own files, permissions, and network rules. Agents never touch your internal systems. Per-agent file permissions. Network egress protection built in.
- Stream: Every event, live in your app. Tool calls, file writes, and progress stream back as they happen. Your users watch the agent work. Server-sent events, no polling. OpenTelemetry traces on every service.
- Measure: Every run priced to the cent. Every run is priced the moment it finishes. Spend caps stop runaway agents before the bill. Per-run cost attribution. Fail-closed monthly spend caps.
Features
Queues, sandboxes, spend controls, observability. One platform, not five side projects.
Orchestrate: Run many agents as one system
Runs, sessions, pools, and workflows as API calls. One integration for Claude, Codex, and Vercel AI.
- Ship agent features in a sprint, not a quarter. Queuing, retries, and run state are handled before you write a line. (Managed run lifecycle, Agent profiles, Persistent sessions)
- Never get locked into one agent SDK. Claude, Codex, and Vercel AI behind one API. Switching is a config change. (Multi-SDK agents, Capability routing, MCP servers)
- Put teams of agents on one problem. Fan a job out across a pool and collect the results deterministically. (Agent pools, Workflows, Skills, Memory bank)
Workspace: Real files, without the risk
Agents get a real filesystem. You keep control of every byte.
- Real tools, zero reach into your systems. Agents read, write, and execute inside an isolated sandbox. (Virtual filesystem, Sandboxed execution, Policy-gated commands)
- Permissions that act, not alert. Access is enforced the moment a file opens. Secrets stay out of prompts and logs. (Per-agent permissions, Scoped secrets, Environments)
- Collaboration without side channels. Pooled agents hand work to each other through files you can inspect. (Pool-shared storage, Role-aware partitions)
Govern: Spend, access, and egress under control
Who can run what. How much it spends. Where its traffic goes.
- Runaway agents stop at the gate, not on the invoice. Spend caps fail closed. When the budget is gone, new runs are refused. (Fail-closed spend caps, Rate limiting)
- Walk into the security review prepared. RBAC, audit log, and network rules in the core platform. (Organizations and RBAC, Audit log, Egress guard)
- A bill that traces back to specific runs. Every run is metered and priced. No end-of-month archaeology. (Usage metering, Credit wallet billing)
Observe: End-to-end observability, by default
Follow any run live, event by event. No black boxes.
- Debug runs while they execute. Every tool call and token streams back over SSE the moment it happens. (Live run streams, SSE into your app)
- Cost per feature without a spreadsheet. Usage rolls up per organization, per agent, and per run. (Usage dashboards, Per-run cost attribution)
- Your observability stack, not another one. OpenTelemetry traces and logs into any OTLP sink you already run. (OpenTelemetry built in, Any OTLP sink)
Integrations
Your backend calls Orca with the OpenAI client. Live agent runs come back.
- API. The OpenAI Agents API: agents, sessions, sandboxes, streaming and files.
- SDKs. No new SDK. The openai package for Python and TypeScript works as it is.
- MCP. Run orca mcp serve and Claude Code, Cursor or Codex build agents for you.
API reference: https://docs.orcapods.ai/reference/overview
Deploy your way
Start in minutes on Orca Cloud, or run Orca’s sandboxes on your own infrastructure. Your agent API stays the same.
We run it. You ship.
Launch on Orca Cloud with nothing to operate. Build, run and scale from one place.
- Live in minutes
- Machine time included in every plan
- Scaling and upgrades handled for you
Your infrastructure. Your terms.
Keep the work on your own hardware: Orca’s sandboxes on your servers or bare metal, on terms shaped around your team.
- Sandboxes on your servers or bare metal
- Dedicated capacity and negotiated rates
- SSO, role-based access and an SLA
FAQ
Direct answers to the questions engineering teams ask first.
Is this only for coding agents?
No. An agent can be a support triager, a research analyst, an operations bot or a code reviewer. Orca runs them all the same way: a model, instructions, tools and skills.
Why not just call the model APIs directly?
You can, until you need sessions that survive a dropped connection, a sandbox per run, spend limits that hold and files you can download. Orca is the layer you would otherwise spend months building.
Do I need a new SDK?
No. Orca speaks the OpenAI Agents API. Point the openai package you already use at https://api.orcapods.ai/v1 and keep your code standard.
How does it fit into our existing app?
Like any backend dependency. Your app starts a session with one call, Orca runs the agent to the end, and the events stream back into your own UI.
What does it cost to run?
A plan includes machine time. On Orca credit, models cost what OpenRouter charges plus its 5.5% fee, with no Orca markup. On your own provider key, Orca charges nothing for tokens. Details on the pricing page.
Which models can I use?
Any model in OpenRouter’s catalog on Orca credit, or your own key for OpenAI, Anthropic, OpenRouter, Vercel or cheaperinference. The model name says who pays, and Orca never switches it for you.
How do you keep agents from doing damage?
Each session works in its own sandbox. Every paid call is checked against your balance first, and a cap on calls in flight stops a runaway agent.
How do we see what agents are doing?
Live. Every message, tool call and file streams back while the agent works, and a dropped stream picks up where it left off.
Can we run it on our own infrastructure?
Yes. We can run Orca’s sandboxes on your own servers or bare metal, with dedicated capacity, negotiated rates and an SLA. Talk to us and we reply within one business day.
Built by you, or your agents
Paste one prompt into Claude Code, Cursor, or Codex. It installs the CLI, signs in with a device code, and follows the guide to build your first agent.
Onboard me to Orca (orcapods.ai), the cloud platform that runs AI agents through the OpenAI Agents API.
1. Install the Orca CLI: `curl -fsSL https://orcapods.ai/install.sh | sh` (https://docs.orcapods.ai/cli/install).
2. Run `orca login`. It prints a one-time code and a URL: relay both so I can approve the sign-in on another device.
3. Add Orca as an MCP server for yourself, as https://docs.orcapods.ai/cli/mcp shows. In Claude Code: `claude mcp add orca -- orca mcp serve`.
4. Ask me what job I want done before you build anything. Then create the agent: a model, instructions, tools and skills. The model name says who pays: `orca/openrouter/<id>` runs on Orca credit, `<provider>/<id>` on my own key for that provider.
The docs are at https://docs.orcapods.ai; start with the quickstart.
Deploy now · Read the AI assistant guide