Surplus Compute

Efficient AI

Choose when your AI work runs.

Now costs more. Later costs less.

POST /v1/batches
{
  "model": "glm-5.2",
  "completion_window": "6h",
  "messages": [
    {
      "role": "user",
      "content": "Summarize these 40,000 support tickets."
    }
  ]
}

A drop-in API for leading open models. One new field: your deadline.

Time turns into efficiency.

Your timeline gives us room to optimize. Surplus hosts models to maximize high-quality throughput and passes the savings on. Same models, same performance. Pay for work, not gaps.

Run now~55% utilized
Immediate execution leaves gaps
Given time~98% utilized
Time lets the same work pack tight

Efficiency means savings.

The more time you give, the less you pay. Just say when.

today's market surplus
Surplus reaches batch-level pricing by 6 hours, underneath today's market line price now 6h 12h 18h 24h on-demand batch · 24h surplus

Join the beta.

We're onboarding a small group of teams & developers now.