SAVED 45 MIN  ·  ▸ READ 3.9 MIN

The AI Inference Supercycle: Building Infrastructure for a Billionx Scale

Stanford Online · 49:15 runtime

Inference demand will grow a billionfold, and the sustainable path for AI companies is to shift from expensive frontier models to post-trained open-source models running on optimized, multi-cloud infrastructure like Base 10's platform.

▸ 0:31

Founding Story: From Finance to AI Infrastructure

A Non-Linear Path to AI Infrastructure

Tuhin’s career began in investment banking, but boredom led him back to engineering and machine learning research. His early startup experiences eventually converged on building Base 10 to support the explosive growth of AI.

Career Journey

2011
Investment banking at Macquarie, privatizing toll roads and airports.
2012
Moves to Boston for machine learning research on neuromuscular disease.
2013
Relocates to San Francisco to join early-stage tech startups.
2015
Starts first couple of companies, but they go nowhere.
2019
Co-founds Base 10 with Phil and Amir to build production inference infrastructure.
Let's build an infrastructure business alongside it ... production inference.
▸ 3:20

Customer Spotlight: Powering Voice AI and Healthcare Scribes

Full-Stack AI Infrastructure

Base 10 runs all optimizations and infrastructure for its customers, managing multiple models to ensure low latency and high reliability, so companies can focus on their core product.

Models Managed per Customer

Whisper Flow (voice typing)
3–4 language models + 2 audio models
Abridge (healthcare scribe)
Pipeline from speech-to-text to clinical note
We run all the optimizations, we run all the infrastructure so that latency from the time you talk to when text shows up is as quick as possible.
▸ 6:16

Why Base 10 Over Hyperscalers: Performance, Reliability, and Developer Experience

The Base 10 Thesis

While 95% of inference spend goes to frontier models, the path to profitable, defensible AI companies lies in custom models—and Base 10 aims to be the platform where they run.

95%
of inference spend goes to frontier models, leaving 5% for custom models, where Base 10 sees the real moat.

The Base 10 Migration Path

  1. Companies first try hyperscalers (AWS, GCP, Azure) for inference.
  2. They struggle with DIY optimizations, reliability, and tooling.
  3. They move to Base 10 for an integrated, fault-tolerant platform.
they realize they'll just be better served coming to base 10.
▸ 8:41

The Economics of Post-Trained Open Source Models

The dual rationale for post-training open source models

Viable: Companies scaling from product-market fit need to boost gross margins from near 0% to 40–70% by running open source models 70–90% cheaper. Cynical: Using frontier models risks giving away proprietary data and user signals, allowing frontier labs to eventually compete with your core workflows.

90 days
Open source models trail frontier models by roughly one quarter.

How to own your intelligence

  1. Adopt open source models that are ~90 days behind but 70–90% cheaper
  2. Post-train on proprietary data to create specialized models for your workflows
  3. Drive gross margins from near zero to 40–70% while maintaining quality
  4. Stay independent from frontier labs to protect your unique user signals
70–90%
Cost savings when running open source models compared to frontier alternatives.
You need to own your intelligence.
Post-trained OSS models deliver frontier-level performance at 30% of the cost, with inference demands surging at the application layer and open-source converging with closed-source within 90 days.Performance vs. cost curveOSS and closed-source convergencePost-training advantages
▸ 9:49Post-trained OSS models deliver frontier-level performance at 30% of the cost, with inference demands surging at the application layer and open-source converging with closed-source within 90 days.
▸ 11:53

Scaling AI Businesses: When to Adopt Post-Trained Models

Scale Drives Adoption of Post-Trained Models

Larger user bases make the cost of frontier model inference unsustainable, making the shift to cheaper, post-trained models an existential priority for business viability.

The larger you are, the more existential it becomes to shift that token volume towards open source.

Post-Training Unlocks Latency and Reliability Gains

By fine-tuning models on their own signals, companies achieve greater control, leading to improved latency, reliability, and overall user experience.

▸ 14:28

Post-Training Workflow with Base 10: From Data to Deployment

The Base 10 Post-Training Promise

Customers define their optimization goal, supply data, and select a base open-source model. Base 10 provides the scaffolding to create a specialized model and seamlessly integrates it into inference, abstracting away all complexity.

Post-Training Workflow

  1. Define your utility function: the metric you want to optimize (e.g., minimize transcription errors).
  2. Provide a dataset relevant to your use case.
  3. Choose a base open-source model (e.g., Kimmy K25).
  4. Base 10 supplies the scaffolding to post-train a specialized model.
  5. The model is automatically integrated into Base 10’s inference stack for deployment.
They come with data and what they know about their workflows, and they leave with a post-trained specialized model running on Base.
The Baseten Platform provides a complete stack for model training, deployment, and inference optimization across multiple cloud providers with automatic GPU capacity management.Multi-cloud capacity managerBaseten Inference StackModel API and Training
▸ 18:01The Baseten Platform provides a complete stack for model training, deployment, and inference optimization across multiple cloud providers with automatic GPU capacity management.
▸ 18:28

Trust, Security, and the Open Source Imperative

Trust Through Strict Security

Base 10 maintains an incredibly intense security posture, setting up internal boundaries to prevent data leakage between competitors, thus earning customer trust.

Open Source as National Security

Open source models are a matter of U.S. national security; the best currently come from China, and if intelligence is 70–90% cheaper in the East, that’s a bad outcome for America.

70–90%
Estimated cost advantage of AI intelligence in Eastern markets vs. the West

Anthropic’s U.S.–China AI Scenarios

U.S. Leads
Shut it down to prevent conflict
Neck and Neck
Leads to war
We think intelligence shouldn't be owned by two people.
▸ 25:40

Hardware Landscape: NVIDIA Dominance and the Rise of Heterogeneous Chips

NVIDIA Dominance in Inference

The majority of inference workloads run on NVIDIA GPUs, propelled by a mature supply chain, strong TSMC relationship, and low cost of capital, making it extremely difficult for competitors to match scale today.

$20B
Amazon's Trainium chip reached a $20 billion revenue run rate, showing the scale of investment in alternative inference hardware.

The Shift to Heterogeneous Architectures

New chips are separating the two core parts of inference—prefill (compute-bound) and decode (memory-bound)—onto different specialized hardware, moving away from the traditional single-chip approach.

there's nothing like CUDA like CUDA is insane.
▸ 29:33

Compute Scarcity and the Rent-vs-Own Pivot

Stitching Compute from 20+ Clouds

Base 10 pools GPUs from over 20 cloud providers and 87 clusters, making compute fungible. This ensures access to scarce hardware by abstracting infrastructure complexity from customers.

30 trillion
Tokens processed daily via Base 10 inference, surpassing the size of OpenAI API and Gemini combined.

Renting vs. Owning Economics

Rental Pricing
Cloud quotes $510/hr (was $263/hr), subject to scarcity pricing
Owned Infrastructure
~30% cheaper at scale, guarantees supply for $7B future demand
$7 billion
Annual compute cost to run 150,000 B200 equivalents, required in 2 years.

Why Supply Won't Normalize

Unlike JFK Airport, AI inference has no daily reset. Demand compounds as models grow and apps become agentic, making GPU access a permanent strategic bottleneck.

▸ 39:19

Future Bets, Student Advice, and Q&A

Next Big Bet: Modular Data Centers

Tuhin would invest in energy and build modular data centers to standardize the unit of compute, similar to how shipping containers revolutionized trade. This would create an API for compute and spur an entire industry.

Advice for Students

Study what excites you; expertise can be built in months. A key emerging skill is project financing for compute infrastructure, given the massive buildout underway.

If the thesis is AGI is everything, every dollar should go to pre-training.