Cloud & Infrastructure

Cerebras and Gimlet Cloud Target 3,000 Tokens/s: What the 100MW AI Inference Deal Means

Gimlet Labs plans to bring Cerebras wafer-scale compute into its multisilicon inference cloud. The headline numbers are ambitious; the more useful question is what they mean for real AI agents, latency, power and deployment timelines.

Digital Pulse Brief  •  Published September 29, 2026  •  Research-based analysis; no hands-on testing claimed

Cerebras Wafer-Scale Engine processor shown in an official Cerebras press image
Image credit: Cerebras Systems — official press kit.

Key takeaways

  • Gimlet and Cerebras say they plan to deploy 100MW of Cerebras-powered inference capacity.
  • The first Cerebras-powered Gimlet Cloud data center is expected later in 2026, while direct CS-4 access for Gimlet customers is expected in 2027.
  • The companies are targeting up to 3,000 output tokens per second for demanding agentic and real-time workloads.
  • That speed figure is a company target/claim, not a Digital Pulse Brief benchmark and not yet a universal production result across models, context lengths and concurrency levels.
  • The technical idea is heterogeneous inference: use different silicon architectures for the stages they handle best rather than forcing every phase onto one accelerator type.

What Cerebras and Gimlet announced

On September 28, 2026, Gimlet Labs and Cerebras announced a strategic partnership to add Cerebras wafer-scale systems to Gimlet Cloud, an inference platform designed to orchestrate AI workloads across different types of accelerators. Gimlet says the companies plan to deploy roughly 100 megawatts of Cerebras-powered inference capacity and target speeds of up to 3,000 tokens per second for agentic and real-time applications.

The two companies also describe a staged rollout. Gimlet says the first Cerebras-powered data center should come online later in 2026, while Gimlet Cloud customers are expected to gain direct access to the next-generation Cerebras CS-4 in 2027. Reuters reported that Cerebras expects to supply CS-4 systems over one to two years and that the financial terms of the deal were not disclosed.

This is best read as infrastructure news with a product consequence: if the promised latency improvements hold under production workloads, developers could give agents more reasoning steps, verification passes and tool calls without making users wait as long.

Why 3,000 tokens per second matters — and why it is not the whole benchmark

Tokens per second measures how quickly a model emits output after generation begins. For a single chat response, faster token generation mostly feels like a smoother interface. For an AI agent that may call a model dozens of times in sequence, latency compounds.

Gimlet illustrates the point with an intentionally simplified example: a multi-step task that takes 10 minutes at 100 tokens per second could, under idealized linear scaling, fall to about 20 seconds at 3,000 tokens per second. That example is useful for intuition, but it should not be treated as a guaranteed end-to-end result. Real workflows also include prompt processing, retrieval, tool execution, network latency, orchestration overhead, context growth, retries and model-specific limits.

For buyers, the better evaluation set is broader: time to first token, output tokens per second, p95 latency, throughput under concurrency, context length, model quality, reliability and cost per useful task. A platform can look spectacular on one speed metric and still be a poor fit for a production workload.

Cerebras AI system hardware shown in an official Cerebras press image
Image credit: Cerebras Systems — official press kit. Hardware shown is an official Cerebras system press asset; it is not presented as a photograph of the Gimlet deployment.

What Cerebras contributes: wafer-scale inference

Cerebras takes a different approach from conventional multi-GPU systems. Its core idea is to put a very large amount of compute and on-chip SRAM on a wafer-scale processor, reducing the amount of data movement and inter-chip coordination that can slow token generation.

The current CS-4 is a rack-scale system built around three WSE-3 Turbo processors. According to Cerebras, a CS-4 provides 750 PFLOPS of AI compute, 129.6 petabytes per second of memory bandwidth and 7.2 terabits per second of I/O, with wafer-to-wafer latency as low as two microseconds. Each WSE-3 Turbo is specified at four trillion transistors, 900,000 AI-optimized cores and 44GB of on-wafer SRAM.

Cerebras also markets CS-4 as delivering up to 30× faster inference than GPU systems and up to 10× more throughput per watt than CS-3. Those are Cerebras performance claims, and the company itself notes that observed speed varies by workload, model, configuration and benchmark method. They should be validated against the exact model and serving conditions a buyer plans to use.

CS-4 at a glance

MetricCerebras-published figureHow to interpret it
AI compute750 PFLOPSPeak system capability; not the same as app-level speed.
Memory bandwidth129.6 PB/sImportant for memory-heavy token generation.
I/O7.2 Tb/sRelevant when linking systems and disaggregating inference.
Wafer-to-wafer latencyAs low as 2 μsHardware-level latency, not full application response time.
Claimed inference advantageUp to 30× vs GPU systemsVendor claim; workload and comparison setup matter.

What Gimlet contributes: a multisilicon inference cloud

Gimlet’s role is not simply to host Cerebras hardware. Its pitch is orchestration across heterogeneous accelerators. Large-model inference contains phases with different compute and memory characteristics, including prefill, token generation or decode, attention, feed-forward work and speculative decoding.

Instead of assuming one processor architecture is optimal for every stage, Gimlet says its software can decompose the workload and map different phases to the silicon that fits them best. The company describes techniques such as prefill-decode disaggregation, attention-FFN disaggregation and speculative-decoding disaggregation, while presenting developers with a standard inference API.

Gimlet claims its heterogeneous approach can produce 3–10× higher interactivity at a given throughput-efficiency target, or a similar gain in throughput per kilowatt at a fixed interactivity target. Again, these are company figures. The important architectural point is broader: the cloud is becoming an orchestration layer for multiple accelerator types rather than a synonym for a homogeneous GPU fleet.

Cerebras wafer-scale computing cluster shown in an official Cerebras press image
Image credit: Cerebras Systems — official press kit. Representative Cerebras cluster image.

What does a 100MW deployment actually mean?

One hundred megawatts is data-center-scale electrical capacity, not the power draw of a single rack or chip. Gimlet says the planned Cerebras-powered capacity will be deployed over time, while Reuters reported a one-to-two-year supply horizon for CS-4 systems.

The exact physical footprint, number of systems, utilization rate, cooling design, energy source and cost structure have not been disclosed in the sources reviewed for this article. Those details matter because the economics of AI inference depend on more than accelerator speed: power delivery, cooling, networking, software efficiency and utilization determine what a provider can actually sell to customers.

The timeline language also suggests a staged build. Gimlet expects its first Cerebras-powered data center later in 2026, but direct CS-4 access for cloud customers is expected in 2027. The companies have not publicly detailed the exact hardware mix at each stage.

What is verified, and what still needs proof

ItemStatusDPB assessment
Cerebras–Gimlet partnershipConfirmedAnnounced by both companies on Sept. 28, 2026.
100MW planned capacityConfirmed planA deployment target, not installed capacity today.
Up to 3,000 tokens/sCompany target/claimNeeds workload-specific production validation.
3–10× multisilicon advantageGimlet claimMethodology and workload details determine relevance.
CS-4 customer accessExpected in 2027Not generally available through Gimlet yet.
Commercial termsNot disclosedNo public deal value or customer pricing in reviewed sources.

Who should care about this

AI-agent builders: Agents that chain many model calls are unusually sensitive to inference latency. Faster generation can create room for more verification or tool use inside the same user-facing time budget. For background, see DPB’s AI agent explainer.

Voice and real-time application teams: Conversational systems feel broken when latency interrupts turn-taking. End-to-end response time matters more than a peak token rate, but faster decode can help.

Infrastructure and platform teams: Gimlet’s multisilicon approach is a practical example of the broader shift from a single-accelerator mindset to workload orchestration. DPB’s cloud computing explainer provides the infrastructure context.

Data-center planners: The 100MW plan reinforces how inference is becoming a major infrastructure workload. Our earlier analysis of grid-aware AI data centers explains why power availability and flexibility are increasingly part of AI-system design.

Buyers comparing inference providers: Do not choose from peak tokens-per-second alone. Ask for reproducible benchmarks on your model, context size, prompt/output mix and concurrency, then compare cost, reliability, availability and operational constraints.

What to benchmark before moving a workload

  1. Use the same model or a quality-equivalent model on both platforms.
  2. Measure time to first token and output tokens per second separately.
  3. Test realistic context lengths, not only short demo prompts.
  4. Run production-like concurrency and record p50/p95/p99 latency.
  5. Measure complete agent-task time, including retrieval and tools.
  6. Compare cost per successful task, not only cost per token.
  7. Check reliability, rate limits, geographic availability and data-handling requirements.

Watch: Cerebras explains CS-4 and wafer-scale inference

Source: Cerebras official YouTube channel. The Supernova 2026 keynote includes the CS-4 and WSE-3 Turbo launch plus demonstrations and discussion of low-latency inference.

Bottom line

The Cerebras–Gimlet partnership is notable less because of one headline speed number and more because of the architecture behind it. Cerebras is betting that wafer-scale memory bandwidth can make token generation dramatically faster, while Gimlet is betting that a cloud can route different parts of inference to different types of silicon.

If that combination holds up under real concurrency, long contexts, tool-heavy agents and transparent pricing, it could make low-latency inference a competitive feature rather than a niche benchmark. But the current 3,000-tokens-per-second and 3–10× efficiency/interactivity figures remain vendor claims. The next useful evidence will be reproducible production benchmarks, supported model lists, customer pricing and reliability data as the rollout moves toward CS-4 access in 2027.

Frequently asked questions

Is Gimlet Cloud offering Cerebras CS-4 today?

Not according to the September 28 announcement. Gimlet says the first Cerebras-powered data center is expected later in 2026, while direct CS-4 access for Gimlet Cloud customers is expected in 2027.

Is 3,000 tokens per second independently verified?

The 3,000-tokens-per-second figure is a target/claim from Gimlet and Cerebras in their announcement. Digital Pulse Brief has not independently benchmarked the service, and the reviewed sources do not establish that rate across every model, context length or concurrency level.

Why combine GPUs and wafer-scale accelerators?

Different inference phases have different compute and memory characteristics. Gimlet’s architecture is designed to map each phase to the type of silicon it expects to handle that work most efficiently, while keeping a common API for developers.

What should companies test before switching inference providers?

Use the same model and realistic prompts, then compare first-token latency, output speed, concurrency, p95/p99 latency, full agent-task time, quality, cost, reliability, data controls and regional availability.

DIGITAL PULSE BRIEF NEWSLETTER

Get clear AI, technology and business insights in your inbox

Breaking developments, practical explainers, reviews and useful tech intelligence — without the noise.

You can unsubscribe from future emails at any time.