NVIDIA researchers introduce Physis-Lang, boosting Cosmos3 past Veo 3.1
A joint team from NVIDIA, MIT and the University of Oxford presented Physis-Lang, a framework that treats physical language as an optimizable representation for video models. The system adds a physics‑reasoning field to captions and uses a self‑evolving loop where a frozen GPT‑5.5 captioner is guided by a Gemini‑3.1‑Pro critic, improving caption F1 from 78.64 to 87.82 after nine iterations. On…
Key points
- Physis-Lang framework adds physics‑reasoning captions, reaching 48.2 ± 1.4 on Physics‑IQ Verified with Cosmos3‑Super.
- Cosmos3‑Nano with Physis‑Lang beats Google Veo 3.1 on three of four benchmarks, e.g., 43.41 vs 34.99 on Physics‑IQ Verified.
- Local PhysThinker‑C/U models keep most gains, reducing API expense from $24.12K to $0.12K.
When compared with Google’s Veo 3.1, Cosmos3‑Nano equipped with Physis‑Lang outperformed on three of four benchmarks, for example scoring 43.41 versus 34.99 on Physics‑IQ Verified. The researchers also distilled the pipeline into two local Qwen‑based models, PhysThinker‑C and PhysThinker‑U, which retained most of the gains while cutting API costs from about $24.12K to $0.12K, and a fully local setup cost $0. No code or weights have been released yet; only the paper and repository are available.
NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks
MarkTechPost · 30 September 2026
Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- GPT-6.1 Sol ranks second in Mahjong AI benchmark behind GPT-6 Astra · 1 src
- Mirror-Score benchmarks D-peptide design tools against real-world affinity · 1 src
- Researchers introduce coffee framework for discrete diffusion model guidance · 2 src
- Researchers propose NashEval for context-dependent AI agent evaluation · 2 src
- Researchers test whether AI harnesses specialize or just repeat answers · 1 src
Comments
via GitHub Discussions