July 19th, 2026
Scaling Human Prescience: VLMs, VL-JEPA, and Synthetic
Trajectories at Scale
Language and vision models scaled because the internet had
already done the data collection for them — trillions of
tokens of text, billions of labeled and unlabeled images,
sitting there waiting to be trained on. Robotics never had
that gift…
Read the post [→]