Surplus Compute

Latency-tolerant inference

Work that can wait
should cost less.

Surplus Compute studies how to run AI workloads more cheaply when results are needed in hours.

Research questions

Use cases

What work can wait?

One engineering lead described a team using real-time inference to analyze support tickets overnight. Another applied standard API calls to an LLM-as-a-judge on chatbot traces. We’re studying workloads like these, and work teams can’t yet afford to run.

Discuss your workload

Technology

How much cheaper can we make it?

For example, giving a job six hours to finish creates room to optimize its cost. We’re building and testing serving systems to measure those savings while preserving quality and meeting deadlines.

Collaborate on the research

Application

What would make teams actually use it?

Discounted inference already exists. Tell us where today’s options fall short for your team, whether it’s model quality, reliability, cost, or integration.

Share what’s missing

About us

We started this project because giving AI more time can make it less expensive. We’ve since spoken with teams and learned that the theoretical savings aren’t in production. We’re exploring how to bridge the gap.

Charles Pollnow

Charles initiated DoorDash’s agentic commerce strategy & go-to-market. Previously, he built and led its grocery business and launched new markets across the US, Canada, and Australia.

LinkedIn

Peter Bromley

Peter built production NLP and distributed systems at Primer AI, authentication software at Badge, and conducted machine learning research at the Altius Institute.

LinkedIn