Data as of Aug 25, 2026 · Based on 42 AI responses from ChatGPT Search and Google AI Mode · See how Parse measures this
Your brand can be here too.
The modern AI stack has expanded far beyond foundational frontier labs (like OpenAI, Anthropic, and Google DeepMind) into a dense, highly competitive ecosystem of AI infrastructure startups . These companies build the plumbing, guardrails, compute engines, and data loops that make large-scale models reliable and deployable.
As frontier models and autonomous agents grow more complex, deterministic testing is insufficient. Startups in this layer provide LLM-as-a-judge frameworks, regression tracking, and CI/CD quality gates.
DeepEval ) and enterprise platform providing automated unit testing, multi-turn conversation simulation, and adversarial red-teaming.Frontier labs no longer just need raw web scrapes; they need high-density reasoning data, expert human verification, and synthetic data generation.
Running frontier models efficiently requires aggressive memory-bandwidth optimization, speculative decoding, and dynamic request routing.
While raw context windows are expanding, production agents require standalone long-term memory systems to recall past user interactions, preferences, and enterprise context across sessions.
With enterprise adoption tied directly to compliance, data privacy, and jailbreak prevention, safety infrastructure has shifted from theoretical alignment to active runtime governance.
Would you like to explore how the memory layer integrates with vector databases , or dive deeper into the economics of production inference providers?
Yes. If you mean the picks-and-shovels startups that become more valuable as the number of frontier/model labs increases, I’d map the market roughly like this.
The important shift is that the stack is moving above raw GPUs: new labs can rent compute, but they still need proprietary data, post-training loops, evals, inference optimization, persistent agent state, and safety infrastructure.
| Layer | Startups I’d watch | What they provide |
|---|---|---|
| Data / post-training | Mercor, Scale AI, Surge AI, Turing, Deccan AI | Expert data, RLHF/RLAIF, coding/agent data, human evaluation |
| Evals / observability | Braintrust, Langfuse, Patronus AI, Arize, LangChain/LangSmith | Testing, tracing, regression detection, model/application evaluation |
| Inference | Fireworks AI, Baseten, Together AI, Modal, Replicate, fal.ai, RunPod | Serving models cheaply/quickly, custom deployment, GPU orchestration |
| Model routing | OpenRouter, Vercel AI Gateway | Routing across many model providers; abstraction above individual inference vendors |
| Agent memory | Mem0, Zep, Letta, Supermemory, Cognee | Persistent memory, temporal knowledge, context management |
| Safety / security | Lakera, HiddenLayer, Protect AI, Giskard, Gray Swan AI, Haize Labs | Red teaming, model/app security, jailbreak detection, guardrails |
| Data infrastructure / knowledge | Databricks, Pinecone, Weaviate, LanceDB, MotherDuck | Retrieval, vector/columnar data, data processing and knowledge infrastructure |
The interesting part isn't traditional labeling anymore. It's post-training data and environments.
Mercor is particularly interesting because it has moved toward expert-generated datasets, benchmarks and RL environments rather than simply selling annotation labor. Its own description explicitly positions it around benchmarks, evaluation environments and expert datasets for frontier labs.
I'd put these into two buckets:
Deccan AI is a useful signal: it raised a $25M Series A specifically around post-training data and evaluation work, including coding and agent capabilities.
Thesis: the scarce resource is increasingly not tokens of data but high-information training experiences.
This may be the most strategically important layer.
The frontier labs can build their own internal benchmarks, but every company building on those models needs to answer:
"Did the new model/prompt/agent actually get better?" That's why I'd watch:
Braintrust and Langfuse are an especially interesting pair: one is a polished commercial platform, the other an open-source/self-hostable alternative.
The deeper opportunity: evals become the equivalent of CI/CD for intelligence. Once models and agents are changing every day, you need automated tests before you can safely ship those changes.
This is probably the most mature startup market on the list.
My current tier-one group:
Fireworks, Baseten and Together are particularly notable because they're attacking the economics of running open-weight/custom models rather than simply reselling frontier APIs. Current industry comparisons put those three at the center of the production-inference market.
Then there's a second layer:
Inference routers
These don't necessarily want to own the GPU. They want to own the decision about which model gets each request.
That could become extremely valuable if model choice becomes dynamic:
request → classify → route → infer → evaluate → retry/escalate In that world, the router starts looking more like a model operating system.
This is the one I'd watch most closely if you're looking for new companies rather than established infrastructure.
The core players:
But "memory" actually contains several different businesses:
Personal memory → "What does this user like/know?"
Episodic memory → "What happened in previous sessions?"
Semantic memory → "What facts has this agent learned?"
Temporal/relational memory → "What was true about X at a particular point in time?"
Agent working memory → "What should remain in context while this autonomous process runs?"
Mem0, Zep and Letta are making notably different architectural bets: Mem0 is a general memory layer, Zep/Graphiti emphasizes temporal knowledge graphs, while Letta treats memory as part of the agent runtime itself.
My bet: memory eventually gets absorbed into agent runtimes/models, but there's a very interesting window before that happens.
This is important.
These companies are closer to:
"Can we discover what this model can do wrong?"
Red teaming, adversarial testing, automated jailbreak discovery, capability evaluation, etc.
These are closer to:
"Can we keep an AI application from being compromised?"
HiddenLayer, for example, now positions its platform across AI discovery, supply-chain security, attack simulation and runtime security.
That distinction matters because agent security is becoming cybersecurity: prompt injection, malicious tools, data exfiltration, compromised MCP servers, model supply-chain attacks, excessive permissions, etc.
I'd divide the opportunities into three tiers.
This is where I think it gets particularly interesting:
The common thread is that the new AI labs are increasingly becoming learning systems, not merely model-training organizations. They continuously collect experience → generate data → train/post-train → evaluate → deploy → observe failures → generate more data.
So the really interesting infrastructure company may sit across several of these boxes, rather than cleanly inside one.
FRONTIER / NEW AI LAB
│
┌────────────────┼────────────────┐
│ │ │
DATA MODELS COMPUTE
│ │ │
Mercor/Scale OpenAI/etc. GPU clouds
Surge/Deccan │
│ Inference
▼ Fireworks/Baseten
POST-TRAINING │
│ ▼
├──────────────► EVALS ◄─── Observability
│ │
│ ▼
│ SAFETY
│ Red team / guardrails
│ │
▼ ▼
AGENT RUNTIME
│
┌──────┴──────┐
▼ ▼
MEMORY TOOLS/MCP
│ │
└──────┬──────┘
▼
PRODUCTION
│
└────► new data
│
└──────► training loop
The biggest strategic insight: inference is becoming commoditized infrastructure, while data + evals + feedback + memory + safety form the closed-loop infrastructure around intelligence. That's the part I'd pay the most attention to over the next 2–3 years.
And the market is already moving in that direction: inference platforms are proliferating rapidly, while training-data companies are shifting toward expert data, evaluations and RL environments.
Yes. If you mean the picks-and-shovels companies that become more valuable as frontier-model labs multiply, rather than the labs themselves, I’d map the market roughly like this:
| Layer | Startups I’d watch | What they’re becoming |
|---|---|---|
| Evals / observability | braintrust.dev, arize.com, smith.langchain.com, patronus.ai, giskard.ai | The testing/CI/monitoring layer for models and agents |
| Training / post-training data | scale.com, surgehq.ai, mercor.com, snorkel.ai | The data factories feeding pretraining, SFT, RL and agent training |
| Inference | fireworks.ai, baseten.co, together.ai, modal.com, deepinfra.com, openrouter.ai | The independent serving layer between models and applications |
| Memory / context | letta.com, mem0.ai, getzep.com | Persistent state, long-term memory and context management for agents |
| Safety / security | lakera.ai, patronus.ai, giskard.ai, robustintelligence.com | Red teaming, prompt-injection defense, guardrails and runtime controls |
1. Scale AI — data + evals + post-training
Scale is arguably the clearest example of an AI-lab infrastructure company. Its Data Engine spans data generation, RLHF, evaluation and red teaming, while Scale Labs now explicitly works on agents, post-training, safety, evaluation and alignment.
The interesting evolution is from “labeling company” → training-data operating system for frontier labs. Scale says its data business now supplies training/test/improvement infrastructure to leading AI labs, and its newer frontier-data products include agent and complex-reasoning datasets.
2. Fireworks / Baseten / Together — the independent inference layer
This may be the biggest infrastructure opportunity after GPUs.
The core thesis is that increasingly many labs will produce models, but relatively few application companies will want to operate the serving stack themselves. Fireworks, Baseten and Together are competing to own that layer: GPU orchestration, optimized kernels, quantization, autoscaling, fine-tuning and model APIs. A recent 2026 industry survey identifies these three, along with Modal, as the leading independent inference platforms.
I'd distinguish them:
The OpenRouter thesis is particularly interesting: instead of owning compute, own the traffic and routing layer across model providers.
3. Letta + Mem0 — memory becomes a primitive
This category is earlier but potentially enormous.
letta.com is going after persistent, stateful agents rather than simply adding a vector database. Its current architecture treats memory, skills, computer use and agent execution as one runtime, with memory decoupled from the underlying model provider.
mem0.ai is taking a more infrastructure-like approach: a persistent memory API that sits between the agent and storage, automatically extracting, updating and retrieving memories. It says it has more than 160,000 developers and has raised $24M.
My bet here is that “memory” eventually looks less like a feature and more like a database primitive—just as caching, identity and search became infrastructure.
4. Patronus — evals evolving into simulated worlds
Patronus is one of the more interesting companies because it's moving beyond static benchmark/eval dashboards. Its 2026 strategy includes simulated digital environments where agents can be stress-tested on realistic, multi-step tasks. It raised $50M in June 2026 for that direction.
That's important because frontier-agent evaluation increasingly becomes:
model → environment → trajectory → outcome rather than:
prompt → answer → score. That creates a much larger infrastructure market.
5. Lakera / Giskard — security becomes runtime infrastructure
The security problem changes radically once models can use tools, access databases and take actions.
lakera.ai is positioned around runtime protection, prompt-injection prevention, data-leakage protection and AI red teaming.
giskard.ai is pushing toward continuous, agent-specific red teaming: dynamically generating attacks based on the agent's behavior and business context rather than relying on static test cases.
I think agent security is substantially bigger than “LLM guardrails.” The eventual product is closer to a combination of Cloudflare + application security + runtime authorization for autonomous software.
I'd draw the new AI-lab infrastructure stack like this:
Data → training data, preference data, synthetic data, agent trajectories ↓
Training / post-training → RL environments, reward models, fine-tuning ↓
Models → frontier labs / open-weight labs ↓
Inference → serving, routing, optimization, GPU orchestration ↓
Context → retrieval, memory, state, identity ↓
Agents → tools, computer use, multi-agent execution ↓
Evals → benchmarks, simulations, observability, outcome measurement ↓
Safety → red teaming, runtime security, policy enforcement, monitoring
The really interesting thing is that the boundaries are collapsing. Scale is already spanning data + evals + safety; Patronus spans evals + safety + simulated environments; Letta spans memory + agent runtime; inference companies increasingly incorporate fine-tuning and evaluation.
My highest-conviction categories would be:
And I'd watch Surge AI particularly closely on the data side: it is now selling frontier training data, evals and RL environments as off-the-shelf products, including coding-agent and enterprise-agent datasets.
The deeper investment thesis is: the number of frontier labs may grow 10×, but the infrastructure required to train, evaluate, serve, remember and secure their models may grow even faster. The winners don't necessarily need to build the best model—they need to become unavoidable pieces of the model-development loop.
My bet here is that “memory” eventually looks less like a feature and more like a database primitive—just as caching, identity and search became infrastructure.
4. Patronus — evals evolving into simulated worlds
Patronus is one of the more interesting companies because it's moving beyond static benchmark/eval dashboards. Its 2026 strategy includes simulated digital environments where agents can be stress-tested on realistic, multi-step tasks. It raised $50M in June 2026 for that direction.
That's important because frontier-agent evaluation increasingly becomes:
model → environment → trajectory → outcome rather than:
prompt → answer → score. That creates a much larger infrastructure market.
5. Lakera / Giskard — security becomes runtime infrastructure
The security problem changes radically once models can use tools, access databases and take actions.
lakera.ai is positioned around runtime protection, prompt-injection prevention, data-leakage protection and AI red teaming.
giskard.ai is pushing toward continuous, agent-specific red teaming: dynamically generating attacks based on the agent's behavior and business context rather than relying on static test cases.
I think agent security is substantially bigger than “LLM guardrails.” The eventual product is closer to a combination of Cloudflare + application security + runtime authorization for autonomous software.
I'd draw the new AI-lab infrastructure stack like this:
Data → training data, preference data, synthetic data, agent trajectories ↓
Training / post-training → RL environments, reward models, fine-tuning ↓
Models → frontier labs / open-weight labs ↓
Inference → serving, routing, optimization, GPU orchestration ↓
Context → retrieval, memory, state, identity ↓
Agents → tools, computer use, multi-agent execution ↓
Evals → benchmarks, simulations, observability, outcome measurement ↓
Safety → red teaming, runtime security, policy enforcement, monitoring
The really interesting thing is that the boundaries are collapsing. Scale is already spanning data + evals + safety; Patronus spans evals + safety + simulated environments; Letta spans memory + agent runtime; inference companies increasingly incorporate fine-tuning and evaluation.
My highest-conviction categories would be:
And I'd watch Surge AI particularly closely on the data side: it is now selling frontier training data, evals and RL environments as off-the-shelf products, including coding-agent and enterprise-agent datasets.
The AI ecosystem has shifted from a pure "frontier model training" race to an intense focus on the supporting infrastructure stack. As frontier labs (OpenAI, Anthropic, Google, Meta, DeepSeek) push capabilities further, a robust ecosystem of specialized startups is building the plumbing for evals, data, inference, memory, and safety.
Inference is where ongoing compute demand lives. Startups here focus on maximizing GPU efficiency, reducing latency, and serving open-source or custom models at scale.
Because LLMs and agents are non-deterministic, debugging them requires entirely new evaluation frameworks to test accuracy, regression, and behavior.
With the world running low on fresh human-generated text for pre-training, data startups focus on cleaning unstructured data and generating high-fidelity synthetic data.
Expanding context windows is hardware-heavy, so memory infrastructure relies on efficient vector databases and state orchestration layers to give models long-term memory.
As enterprises deploy agents that can take actions, safety has morphed from theoretical alignment into enterprise compliance firewalls, risk platforms, and runtime guardrails.
If you want to zoom in on one of these layers, let me know which area (Inference, Evals, Data, Memory , or Safety ) you want to explore, and I can break down their technical architectures or market positioning.
The AI infrastructure stack has evolved past raw GPU compute into a specialized ecosystem of software and hardware control points. Startups are tackling major bottlenecks across evals, data, inference, memory, and safety:
Would you like to explore the economic and pricing models of inference vs. training infrastructure , or dive deeper into one specific layer like agent memory architecture?
Modal
Weka
weka.io
If you mean the startups that become indispensable suppliers to frontier-model labs, rather than generic “AI infrastructure,” I’d map the market roughly like this as of August 2026:
These are building the measurement and feedback loop around models and agents.
The interesting shift is that evals are becoming part of the training loop, rather than simply a QA product. The emerging stack is: generate data → train/post-train → evaluate → identify failures → generate more data → repeat. Current market surveys put Braintrust, Arize, Langfuse, Patronus and others squarely in this layer.
This may be the most strategically important layer for frontier labs.
The really interesting subcategory here is RL environments. Mercor, for example, explicitly describes building realistic environments, tools, tasks and verifiers for frontier models.
That suggests the next generation isn't merely “label 10 million examples.” It's “construct an environment in which an agent can attempt millions of difficult tasks and automatically receive a meaningful reward.”
This is arguably the hottest infrastructure market right now because inference is becoming the economic bottleneck.
There's a useful structural distinction:
GPU/cloud layer → inference engine → inference platform → model router.
Fireworks/Baseten/Together/Modal are fighting around the middle; OpenRouter is moving above them. Recent market research identifies those companies as the core independent inference platforms, while inference spending itself is becoming a larger share of AI infrastructure economics.
And there is a new hardware layer worth watching: Etched. It just raised $700M at a reported $21B valuation, explicitly betting that inference economics justify purpose-built silicon.
This is the layer I think is most underappreciated.
The architectural battle is fascinating:
vector memory (Mem0) vs. temporal graphs (Zep) vs. agent-managed/paged memory (Letta) vs. knowledge graphs (Cognee).
The common thesis is that a 1M-token context window isn't the same thing as memory. Agents need to store, retrieve, update, reconcile and forget information over months of interaction. Current comparisons consistently put Mem0, Zep and Letta among the leading dedicated systems.
I'd split this into runtime safety, model security, and red teaming.
The category is moving beyond traditional “content moderation.” The emerging problem is agent security: prompt injection, tool abuse, malicious MCP servers, exfiltration, model supply-chain attacks and agents escaping their intended sandbox. Recent industry analysis groups Protect AI, HiddenLayer, Lakera, CalypsoAI and Adversa around this broader AI-security stack.
If I were trying to understand where the next infrastructure layer around frontier labs is forming, rather than compiling every vendor, I'd watch:
| Layer | Companies I'd watch most closely |
|---|---|
| Data / RL | Mercor, Surge, Scale, Handshake |
| Evals | Braintrust, Patronus, Arize |
| Inference | Fireworks, Baseten, Together, Modal |
| Memory | Mem0, Zep, Letta |
| Safety | Lakera, Protect AI, HiddenLayer |
But there's a bigger pattern here.
The most valuable companies may not remain cleanly in one category.
The frontier-lab infrastructure loop increasingly looks like:
Data → post-training/RL → evals → inference → telemetry → memory → safety → more data
And the companies with the strongest strategic position are the ones that sit on multiple edges of that loop.
For example, Mercor is moving from expert labor → data → evaluations → RL environments. Fireworks is moving from inference → fine-tuning → evaluation. Memory companies are moving toward agent runtimes. Safety vendors are moving from static guardrails toward runtime agent security.
That's why I'd pay particular attention to Mercor, Fireworks, Baseten, Braintrust, Mem0, Letta and Lakera. They aren't merely selling another developer tool; they're trying to own a piece of the production feedback loop around increasingly autonomous models.
If your underlying question is “which of these could become the next $10B+ infrastructure companies?”, I'd rank the categories quite differently: inference and frontier data are the obvious current winners, while memory and evals may have the more interesting new-platform upside.
The AI infrastructure stack has evolved past raw compute into a sophisticated ecosystem of control points . As inference workflows, agentic systems, and frontier models scale, specialized startups have emerged around key bottlenecks: evaluations, data, inference, memory, and safety/security.
Evaluating non-deterministic model outputs, scoring agent behaviors, and regression testing require dedicated evaluation infrastructure:
Managing unstructured data pipelines, high-value domain data, and reinforcement learning feedback loops remains strategic:
With inference accounting for the vast majority of active AI compute, serverless execution and intelligent routing are dominant:
Overcoming the "memory wall" and bandwidth limitations for larger context windows and real-time agents has sparked specialized architecture startups:
As models act autonomously via agents, security has expanded to combat prompt injection, data exfiltration, and model abuse:
If you'd like to dive deeper, let me know which of these layers (Evals, Data, Inference, Memory, or Safety ) you want to explore further, or if you need help mapping a specific startup's funding or technical architecture.
Yes. If you think of the frontier-model lab as the new “application,” there’s a pretty interesting picks-and-shovels layer forming around it. The important shift in 2026 is that the stack is moving beyond “LLM APIs + vector DBs” toward post-training data, agent environments, inference economics, persistent state, and continuous safety/evaluation.
Here’s how I’d map the startups I’d be watching.
| Layer | Startups to watch | What they’re really selling |
|---|---|---|
| Evals / observability | Braintrust, Patronus AI, Arize, Galileo, Langfuse | The CI/CD, testing, tracing and quality-control layer for models and agents |
| Training / post-training data | Scale AI, Surge AI, Mercor, Gretel | Human expertise, preference data, synthetic data and increasingly RL environments |
| Inference | Baseten, Fireworks AI, Together AI, Modal, fal, OpenRouter | Turning expensive model weights into reliable, cheap, low-latency production compute |
| Agent memory / state | Letta, Zep, Mem0, Supermemory, Cognee | Persistent state, user/agent memory, temporal knowledge and cross-session context |
| Safety / AI security | Lakera, Protect AI, HiddenLayer, CalypsoAI | Red-teaming, model security, prompt-injection defense, runtime controls and AI supply-chain security |
1. Mercor — data is becoming an interface to the real world.
This is probably the most interesting evolution in the data category. Mercor isn't just supplying annotators; it describes its business as building benchmarks, human datasets and RL environments for frontier labs. Its environments increasingly let agents interact with realistic tools and workflows rather than merely answering static questions.
That matters because the bottleneck for the next generation of models may be less “more internet tokens” and more “more high-quality experience of doing difficult things.”
2. Braintrust — the “GitHub Actions for AI quality” thesis.
Braintrust is moving toward a full development loop: traces → datasets → evals → regression tests → production monitoring. Its pitch is essentially that evals become infrastructure rather than a research exercise.
I think this category becomes particularly important for frontier labs selling models to developers: every model release creates an enormous downstream evaluation problem.
3. Patronus AI — potentially more interesting than a conventional eval SaaS.
Patronus has expanded from LLM evaluation into agent evaluation, synthetic environments and world models. Its June 2026 Series B announcement explicitly positioned the company around simulating intelligence and training/evaluating agents.
That makes it a good example of where the stack is going: evals → environments → training data → oversight.
4. Baseten / Fireworks / Together — inference is becoming its own cloud.
This is one of the clearest infrastructure races. Baseten, Fireworks and Together are effectively competing to become the AWS/GCP layer specifically optimized for model inference. Recent industry comparisons put them at the center of the production-inference market, alongside Modal, Replicate, OpenRouter and others.
The interesting part isn't simply GPU access. It's optimization of batching, quantization, scheduling, caching, model compilation, cold starts and routing—all of which can dramatically change inference gross margins.
5. Letta / Zep / Mem0 — memory is trying to become a primitive.
The emerging memory companies are making different architectural bets:
A 2026 comparison identifies essentially this group as the core independent memory market.
The bigger question is whether memory becomes an independent infrastructure layer—or gets swallowed by the model providers' agent runtimes.
6. Lakera / HiddenLayer / Protect AI — safety is converging with security.
This is another category I'd watch closely. Lakera is focused heavily on prompt injection, malicious context and runtime protection; HiddenLayer covers AI discovery, supply-chain security, attack simulation and runtime defense; Protect AI spans model selection/testing through runtime.
That's important because agentic systems blur the boundary between “AI safety” and “cybersecurity.” An agent that can browse, retrieve documents, call tools and execute code has an attack surface much closer to an operating system than a chatbot.
Scale AI is unusual because it touches data + human evaluation + benchmarks + safety + post-training. Its Scale Labs initiative now explicitly covers agents, post-training, reasoning, safety, evaluation and data science.
That makes Scale an interesting bellwether for where the infrastructure market is going: the boundaries between data provider, evaluator and training environment are disappearing.
If I were trying to understand where the next $100B of AI infrastructure value might accrue, I'd start with:
Data / experience
Evals / oversight
Inference
Memory
Safety / security
And there's a sixth category I'd add that you didn't mention: RL environments / agent training infrastructure. That's potentially the most important emerging layer. The thesis is that once frontier labs exhaust the easy gains from static web data, they need environments where models can act, fail, receive rewards, use tools, and learn from trajectories. Current activity around Mercor, Patronus and startups such as Mechanize points strongly in that direction.
The meta-thesis: the winning infrastructure companies may not be the ones selling compute to the labs. They may be the ones selling the labs the feedback loop: generate experience → collect data → train → evaluate → deploy → observe failures → generate better data → retrain.
That's a much bigger market than “LLM observability” as the category was defined two years ago.
The AI infrastructure ecosystem supporting frontier labs and enterprise deployments has shifted dramatically. With inference spending overtaking training in the cloud and the rise of autonomous agent architectures, specialized startups have emerged across the critical layers of the stack.
Evaluating probabilistic models, tracking behavioral drift, and running automated red-teaming have become vital platform layers.
High-cost reasoning models require clean, specialized, and synthetic data to push past the limits of public web scraping.
Maximizing throughput, lowering time-to-first-token, and dealing with specialized silicon constraints are the core focus here.
Stateless foundation models face a hard wall with multi-turn conversations and long-running workflows, prompting a dedicated memory layer.
As agents gain tool-use capabilities and interact with internal corporate networks, runtime security and guardrails have become crucial.
If you're exploring a specific layer, let me know if you would like to dive deeper into the economics of inference vs. training or look closer at the security stack for multi-agent systems.
Mem0
mem0.ai
Zep
zep.com
Scale
scale.com
Mercor
mercor.com
Mem0
mem0.ai
Letta
letta.com
If you mean the picks-and-shovels companies that sit directly around frontier-model labs—supplying the data, evals, inference stack, memory, and safety machinery—I’d map the 2026 landscape roughly like this.
The important shift is that the stack is moving upward from GPUs toward the “model production system”: proprietary data, post-training/RL environments, inference optimization, persistent context, and adversarial testing. Inference in particular is becoming a major economic control point; Gartner estimates inference will account for 55% of AI-optimized IaaS spending in 2026.
Most interesting:
Thesis: the interesting company isn't necessarily another benchmark provider. It's the continuous evaluation system of record: every model version, prompt, agent trajectory, human judgment, regression and production failure flows through it.
This may be the most strategically important category.
There's a fascinating structural point here: the training-data market appears to be concentrating around Scale, Surge, Mercor and Handshake, with estimates putting their combined share above 75%.
My bet: the next really valuable layer isn't generic labeling. It's proprietary environments that generate training signal—simulated workplaces, software environments, scientific environments, games, robots, browsers, etc. Whoever owns the environment can generate essentially unlimited fresh post-training data.
This is probably the hottest infrastructure layer right now.
The emerging architecture looks like:
GPU → inference runtime → model host → routing/gateway → application
rather than every AI company simply calling OpenAI/Anthropic APIs. A 2026 industry survey already identifies Baseten, Together, Fireworks, Modal, Replicate, OpenRouter and others as distinct layers of this market.
The really interesting startups here are the ones improving tokens/$ and tokens/sec, not merely reselling GPU capacity.
This is earlier, but potentially enormous.
The architectural split is already becoming clear: Mem0 is relatively lightweight fact memory, Letta is agent-managed/tiered memory, and Zep/Graphiti is a temporal knowledge graph.
My contrarian take: vector databases aren't necessarily the winner here. The valuable primitive may become “state”—what an agent knows about a user, task, organization, environment and previous actions, including when each fact was true.
If agents become persistent workers rather than stateless chat sessions, this layer could become as important as databases are to conventional software.
I'd split this into model safety and AI security.
The market is converging on continuous adversarial testing → vulnerability discovery → guardrails → monitoring, rather than a one-time “AI safety review.” Gray Swan is a particularly clean example of this loop.
If I had to reduce the whole thing to a shortlist of companies with potentially fundamental infrastructure positions, I'd start here:
| Layer | Companies I'd watch | Why |
|---|---|---|
| Evals | Braintrust, Patronus, Arize | Become the test/observability layer for AI |
| Data | Mercor, Surge, Scale, Handshake | Own scarce frontier-quality human/RL data |
| Inference | Baseten, Fireworks, Together, Modal | Capture the economics of every inference token |
| Memory | Letta, Mem0, Zep | Become persistent state for agents |
| Safety | Gray Swan, Lakera, HiddenLayer | Security/evals become mandatory as agents gain autonomy |
I think there are three particularly attractive primitives emerging:
1. Data → feedback loops
The winners don't just sell datasets. They create an environment → collect agent behavior → evaluate it → generate better training data → improve the model → repeat.
2. Inference → economics
Inference is becoming the equivalent of the cloud-compute bill for AI. Whoever can make a model 2× cheaper/faster has a very tangible economic moat. The recent Baseten/Fireworks/Together fundraising and infrastructure commitments are evidence of how aggressively capital is moving into this layer.
3. Evals + safety → trust
As models become agents, “does this answer look good?” becomes inadequate. Labs need to know what the system can do, where it fails, whether it can be attacked, and whether a new model release regresses on anything important.
That's why I find Braintrust + Patronus + Gray Swan especially interesting: they sit on the measurement/verification side of the stack, which could become the equivalent of testing, observability and security for the AI-native software era.
And I'd put Mercor in a category of its own: if frontier models increasingly run out of useful internet-scale training data, the company that can manufacture new, expert, high-signal training environments could become one of the most strategically important suppliers to the labs.
Modal
Mem0
mem0.ai
Zep
zep.com
Arize AI
arize.com
Langfuse
langfuse.com
Surge AI
surgehq.ai
Deccan AI
deccan.ai
Inferact
inferact.ai
Cognee
cognee.ai
Witness AI
witness.ai
Mercor
mercor.com