βBirdai
Back to Research
AnalysisFebruary 20266 min read

Why the Hardest Path Compounds

AI is rewriting the value of data. The companies building novel, domain-specific ground truth will own the most defensible positions in tech. We explain why we built Birdai from the data layer up.

There is a pattern playing out across the AI industry that has direct implications for how we think about what we are building at Birdai.

Capital is flooding into the application layer of AI. Copilots, chatbots, workflow wrappers. These products are quick to build and easy to demo. They are also easy to replicate. Every time a new foundation model ships, a generation of application-layer startups watches its technical edge evaporate. Sequoia Capital called this dynamic out directly in their analysis "Generative AI's Act Two": if your product only exists because of a deficiency in today's models, you are on a countdown timer [1].

The companies building durable positions are doing something different. They are building at the data layer. Proprietary collection infrastructure. Domain-specific ground truth. The kind of data that AI models need to learn from but cannot generate on their own.

This is not an abstract observation for us. It is the thesis behind Birdai.

The data bottleneck

The popular narrative says AI's primary constraint is compute or architecture. Better GPUs, larger parameter counts, cleverer attention mechanisms. From a research standpoint, this view is getting harder to defend.

Research from Epoch AI estimates the stock of quality-adjusted public human text at roughly 300 trillion tokens. At current growth rates, this supply will come under increasing pressure through the end of this decade [2]. Foundation models have been trained on what is effectively a lossy snapshot of the open internet. That corpus is vast but shallow. It excludes most of the structured, verified, domain-specific knowledge that would be required to build reliable systems in specialized fields.

The bottleneck is shifting from compute to data. Not just any data. High-quality, verified, primary-source data that captures ground truth in domains where no clean dataset currently exists.

A landmark 2024 paper published in Nature demonstrated what happens when models try to substitute synthetic data for the real thing. When generative models are iteratively trained on their own output, they progressively lose the ability to represent the tails of the original distribution. The researchers termed this "model collapse." Synthetic data has its uses in augmentation, but as a primary training source it degrades model quality irreversibly [3].

Conversely, Microsoft Research showed that a small model trained on carefully curated ground truth data achieved performance competitive with models trained on orders of magnitude more web-scraped noise [4]. Quality beats quantity when the quality is real.

Instruments before discoveries

Freeman Dyson argued in Science that "new directions in science are launched by new tools much more often than by new concepts" [5]. Galileo built a telescope before anyone discovered the moons of Jupiter. X-ray crystallography came before the structure of DNA. The instrument creates access to data that was previously unobservable, and that data restructures the field.

The same dynamic applies to AI. The companies that will matter most are not the ones building applications on top of existing models. They are the ones building the instruments: purpose-built data collection infrastructure that captures high-fidelity information in domains where no clean dataset exists.

This is why we built Birdai from the data layer up.

MEV on Sui as a case study

MEV intelligence on Sui is a clear example of the data layer thesis in practice. The data cannot be ported from other chains.

On Ethereum, MEV infrastructure was built on top of a sequential execution model with a public mempool and single block proposers. The entire Flashbots pipeline assumes these architectural properties. On Solana, Jito built a modified validator client that leverages the leader-based block production schedule. Both approaches are tightly coupled to the architecture of their respective chains.

Sui breaks all of these assumptions. Object-centric state. Parallel execution. DAG-based consensus through Mysticeti. No publicly observable mempool in the Ethereum sense. No single block proposer. The tools that work on Ethereum and Solana have no surface to attach to. You have to build from scratch.

This is what makes the data defensible. It requires building collection infrastructure from first principles for an architecture that has no precedent.

Why domain expertise does not transfer

a16z's framework for evaluating data moats maps directly to what we are building [6]. The moat holds when the data is scarce, proprietary, and structurally difficult to replicate. MEV intelligence on Sui meets all three criteria.

You cannot take a data pipeline built for Ethereum MEV and repurpose it for Sui. The execution model is different. The consensus mechanism is different. The data structures are different. The edge cases are different. Domain expertise does not transfer. That is not a limitation. It is the moat.

Winner-take-most dynamics

MEV infrastructure consolidates. This is not speculation. It is the empirical record.

Flashbots relays approximately 80% of all MEV-Boost blocks on Ethereum [7]. Jito's validator client went from 48% to 94% of Solana network stake in a single year [8]. In both cases, the first team to ship production-grade infrastructure became the default because downstream participants converged on a single coordination layer. Switching costs compound. The data moat widens. The position becomes structural.

We see the same dynamics at work in our domain. Once protocols, searchers, and investors build workflows against a data layer, the switching costs become real. The position becomes structural.

Building from the ground truth up

The AI industry's appetite for high-quality, domain-specific training data is growing faster than the supply. The companies building novel ground truth in hard, specialized domains are producing raw materials that the broader ecosystem needs. Not because they set out to build AI training data, but because the data they produce happens to be exactly what AI systems need to learn about domains that the open internet does not cover.

Blockchain execution data is one of those domains. How value moves through DeFi protocols, who captures it, how auction mechanisms perform, where extraction concentrates. This is precisely the kind of structured, verified, longitudinal data that does not exist in any foundation model's training corpus.

We did not build Birdai as an application on top of an existing data layer. We built the data layer itself. That was the harder path. It took longer. It was harder to explain. But it is the path that compounds.

That is what we build at Birdai.

References

  1. [1] Huang and Grady, "Generative AI's Act Two," Sequoia Capital, 2023: sequoiacap.com
  2. [2] Villalobos et al., "Will we run out of data? Limits of LLM scaling based on human-generated data," Epoch AI / ICML 2024: arXiv:2211.04325
  3. [3] Shumailov et al., "AI models collapse when trained on recursively generated data," Nature, Vol. 631, 2024: Nature
  4. [4] Gunasekar et al., "Textbooks Are All You Need," Microsoft Research, 2023: arXiv:2306.11644
  5. [5] Dyson, "Is Science Mostly Driven by Ideas or by Tools?" Science, Vol. 338, 2012: Science
  6. [6] Casado and Lauten, "The Empty Promise of Data Moats," Andreessen Horowitz, 2019: a16z.com
  7. [7] MEV-Boost Relay Block Percentages: mevboost.pics
  8. [8] Bruder, "Jito: Past, Present, and Future": Substack
Back to Research
Cookies

We use only essential cookies and privacy-respecting, cookieless analytics: no cross-site tracking, no ad pixels. Details in our privacy policy.