Memory Palace

Transformative AI and Current Trajectory

Scaling drivers, capability trends, and time-horizon forecasts for thinking about whether AGI-like systems may arrive sooner than institutions expect.

Scaling drivers, capability trends, and time-horizon forecasts.

Transformative AI and Current Trajectory cover

Theme

This week asks whether current AI progress is best understood as a continuation of scale-driven gains or as something closer to a transition into qualitatively new capability regimes. Frontier systems appear to be improving quickly on tasks that matter, but the central question is whether compute, data, inference, and algorithmic efficiency can keep compounding long enough to produce genuinely transformative systems.

A useful framing for the note is that scaling is not a single smooth track. It is better understood as a sequence of temporary solutions, where each apparent breakthrough tends to shift pressure onto a different bottleneck. That makes the trajectory feel both impressive and unstable at the same time.

Current Trajectory

Recent frontier models seem to be climbing a remarkably steep capability curve, especially on software and reasoning tasks. METR’s time-horizon work is helpful here because it translates capability into something more intuitive than benchmark scores: how long a task is, in human time, that a model can complete with a given reliability.

Figure 1. METR’s time-horizon task suite maps model success rate against estimated human time-to-complete.

Figure 1. METR’s time-horizon task suite maps model success rate against estimated human time-to-complete. I read the 50% success point as a practical reliability boundary: not “the model cannot do longer tasks,” but “the model becomes less dependable as the task stretches.”

The important implication is not just that models can sometimes do difficult things. It is that their reliable task length is increasing fast enough to make longer autonomous work feel increasingly plausible. At the same time, the gap between “can sometimes do it” and “can do it robustly” still matters a lot for real-world deployment.

Figure 2. AI time horizons are increasing across several domains, from coding benchmarks to math, web tasks, and agentic environments.

Figure 2. AI time horizons are increasing across several domains. The useful signal here is the slope: progress is not only showing up as higher benchmark scores, but also as longer tasks that models can complete at a fixed success rate.

But a steep capability curve raises an obvious question: can the inputs that produced it keep compounding? To see why that is uncertain, it helps to separate the drivers underneath the trajectory before looking at where they strain.

Scaling Drivers

The usual story is that model performance improves as training compute rises, data becomes more abundant, and algorithmic efficiency improves. In practice, these drivers are tightly interdependent: more compute can partly offset limited data, while better algorithms can make the same compute more valuable. That is why progress can look smooth even when individual inputs are under stress, and why the real question is often which constraint becomes binding next.

Figure 5. Epoch AI trend indicators across inference prices, compute stock, training compute, software progress, data center scale, and FLOP per dollar.

Figure 5. Scaling is not one curve. Inference prices, chip performance, software efficiency, training compute, data center scale, and total compute stock move at different speeds. This is why the trajectory can look smooth at the frontier while still being driven by many uneven infrastructure trends underneath.

The Data Paradox

Data is the cleanest illustration of that question — and of the framing from the start of the note, that one bottleneck can be relaxed only by making another more important. If frontier labs run out of high-quality human data, synthetic data becomes a plausible workaround. But synthetic data is not free; generating it creates new demand for compute, which then pushes against power, chips, and training efficiency.

Figure 3. Projections of public text stock and dataset sizes used to train notable language models.

Figure 3. Public human text as a scaling input rather than an infinite background resource. The exact year matters less than the shape of the constraint: once models can consume most high-quality public text, the question becomes how much value can still be extracted through filtering, reuse, synthetic data, and overtraining.

This is what makes scaling feel paradoxical. The bottleneck does not disappear; it migrates. Data scarcity may be softened, but the system then faces a sharper compute constraint, and compute itself is ultimately tied to infrastructure, supply chains, and capital intensity.

Figure 4. The synthetic-data workaround shifts pressure from human data toward compute, power, chips, latency, and efficiency.

Figure 4. The synthetic-data paradox. Synthetic data may relax the original human-data bottleneck, but it can also increase demand for training runs, compute, power, chips, and efficiency improvements. In that sense, scaling does not escape scarcity; it relocates it.

Bottlenecks To Watch

The most important near-term constraints are power, chip manufacturing, data, and latency. Power and chips appear to be the most binding; data is the most uncertain, because its effective supply depends on quality, modality, reuse, and synthetic augmentation — and, as the data paradox above shows, easing it tends to push pressure straight back onto compute. The practical upshot is that no single input forecast is enough on its own: the trajectory holds only as long as the currently binding constraint keeps yielding.

Endnotes

I am currently reading Michal Meidan, China: Climate Leader and Villain, and will return soon with notes on China’s Energy Revolution in the Context of the Global Energy Transition. Given that China sits at the center of industrial policy, energy transition, and climate politics, that reading should sharpen the geopolitical dimension of this series.

Core Readings