The bottleneck nobody solved
Language and vision models scaled because the internet had already done the data collection for them — trillions of tokens of text, billions of labeled and unlabeled images, sitting there waiting to be trained on. Robotics never had that gift. A trajectory — a robot's stream of observations and actions as it performs a task — has to be physically enacted, in the physical world, usually by a real robot or a human teleoperator, one episode at a time. Even an aggressive data-collection operation tops out at a small fraction of what a single day of internet crawl gives a language model. This is the actual bottleneck behind every "robots are years behind LLMs" observation: not architecture, not compute, but the raw supply of embodied experience to learn from.
Reaction is not enough
Most deployed robot policies are reactive: given the current observation, output the next action. This works until the world does something the policy hasn't seen enacted from that exact vantage point before, at which point the policy has no mechanism for anticipating what's about to happen — it can only respond after the fact. Humans operate differently. A person catching a dropped glass isn't reacting to the glass hitting the floor; they're running a fast, implicit forward model of where the glass will be a few hundred milliseconds from now and moving to intercept it. That forward model — predicting future world-state before it arrives — is what we mean by prescience, and it's not a philosophical flourish: it's a concrete architectural difference between a policy that maps observation→action and one that maps observation→predicted-future-state→action. The second kind is more sample-efficient (it can reason about states it's never directly seen, by interpolating in a learned latent space) and safer (it can flag that a trajectory leads somewhere bad before committing to it).
Borrowing priors from vision-language models
You don't need to teach a robot policy what a cup is, that cups can hold liquid, or that liquid spills when a cup tips past some angle — a vision-language model pretrained at internet scale already encodes all of this as part of its general world knowledge, learned from images, video, and text describing the physical world. The practical move is to not throw that away: use a VLM's visual and semantic representations as the backbone a robot policy sits on top of, so that what remains to be learned from scarce real trajectories is the control residual — the mapping from "the world looks like this and I want this outcome" to "move the actuator this way" — rather than the entire notion of what objects are and how they behave.
Predicting embeddings, not pixels: VL-JEPA
The naive way to build a predictive world model is to predict future video frames directly — pixel by pixel. This is a bad use of model capacity: most of a video frame's pixels (texture, lighting, background clutter) carry no information relevant to the task, and a model that has to reconstruct them wastes representation power on noise instead of signal, and tends to converge slowly or not at all on the parts that matter.
VL-JEPA (a joint-embedding predictive architecture applied to vision-language inputs) sidesteps this. Instead of a decoder trained to reconstruct future pixels, VL-JEPA trains three pieces jointly: a context encoder that embeds the current observation (and language context, if present), a target encoder that embeds a future observation, and a predictor trained to predict the target encoder's embedding of the future frame from the context encoder's embedding of the present one — never touching raw pixels for the future step at all. Informally, the objective looks like:
L = || predictor(encoder(x_t), a_t) − stopgrad(target_encoder(x_{t+k}))
||²
where x_t is the current observation,
a_t the action taken, and x_{t+k} the
observation k steps later. The stop-gradient on the target
branch (standard in JEPA-family objectives) prevents the
trivial collapse where both encoders learn to output a
constant. What this buys you is a predictive model that only
has to get the task-relevant abstraction of the future right —
not its exact pixel rendering — which is both cheaper to train
and, empirically in the JEPA literature, more transferable.
Cosmos and the synthetic trajectory multiplier
Even with a VLM backbone and a JEPA-style objective that makes better use of every trajectory it sees, the ceiling is still set by how many trajectories exist to train on. This is where NVIDIA Cosmos comes in: Cosmos is a family of world-foundation models trained to generate physically-grounded video and simulation rollouts — not just visually plausible video, but rollouts constrained to respect physics (contact, gravity, rigid-body dynamics) closely enough to be useful as synthetic trajectories. We use Cosmos to multiply the effective size of a trajectory dataset: take a modest set of real robot trajectories, and generate a much larger set of synthetic variations around them — different object placements, lighting, textures, minor physical perturbations (domain randomization) — that stay physically plausible. The resulting corpus is used to pretrain the JEPA world model at a scale real data collection alone couldn't reach, with the real trajectories reserved for fine-tuning and closing the residual sim-to-real gap.
The reason physical grounding matters here, specifically, is that a JEPA-style predictor is only as good as the future-state targets it's trained against. If the synthetic data drifts from real physics, the predictor learns to anticipate a world that behaves differently from the one the robot actually operates in — the model still "looks" confident, but its forward predictions become systematically wrong exactly where physical dynamics matter most.
A scaling law for embodied prescience
LLM scaling laws made a specific, falsifiable claim: loss falls as a predictable power-law function of dataset size (and compute, and parameters), which is what justified spending enormous resources on more data rather than cleverer architectures. We don't yet have a rigorously fitted equivalent for embodied trajectory data — that's an open empirical question, not a solved one — but the shape of the claim we're betting on is directly analogous:
policy_success_gap(N) ≈ C · N^(−α)
where N is effective trajectory count (real + synthetic, appropriately weighted), and the gap to some upper-bound performance shrinks as a power law in N, with some constant α we'd expect to depend on task diversity and how well the synthetic distribution matches reality. This is deliberately presented as a hypothesis we're testing, not a fitted result — the honest version of the scaling story here is that the LLM playbook (more data, same recipe, predictable gains) plausibly applies to embodied prescience once the data bottleneck is addressed, and Cosmos-generated synthetic trajectories are the practical lever for moving N by orders of magnitude when real collection can't.
What ships
Put together, "human prescience" is not a metaphor we're leaving at the marketing layer — it cashes out to a specific technical stack: a VLM-initialized encoder, a JEPA-style predictor trained to forecast future latent states rather than pixels, on a training corpus assembled from real robot trajectories augmented by Cosmos-generated synthetic ones at scale, deployed as the foresight layer underneath a robot's control policy. The bet is that the same lever that took language models from narrow to general — orders of magnitude more training experience, used well — is available to embodied AI too, once you stop trying to collect it all by hand.