TypeSafe ai's Jev beats pokemon Red in under a week
Jev, a decision model from TypeSafe AI, defeated Pokémon Red’s Elite Four and Champion, entering the Hall of Fame on September 23, 2026, after a week of play. The project was announced by Andrew Boyd, founder of Standard Agents Inc., and livestreamed in a browser or terminal with Jev moderating chat.
Key points
- Jev, a decision model, beat Pokémon Red’s Elite Four and Champion in a week
- Claude Opus 5 coached Jev by adjusting options and reducing text by two‑thirds
- Christian Mathiesen’s Frigade run cost about $1–1.70 per 24 hours
Jev selects from a list of options and returns solutions with a confidence figure, but it does not read the screen or produce text or images. Anthropic’s Claude Opus 5 monitored the game log and coached Jev by adjusting options and wording. Opus reduced the text sent to Jev by about two‑thirds and changed the request frequency to every six seconds. Human viewers also sent tips that were incorporated into Jev’s option list. The harness changelog recorded 474 entries, many failures such as walking into Lorelei’s entrance 53 times.
Other runs followed. Christian Mathiesen at Frigade used Jev with a memory‑reading harness, estimating a cost of about $1–1.70 per 24 hours. A separate experiment by stmonty trained a small world model on an RTX 3080 Ti with more than 42,000 frames; it picked a starter in 52 of 100 tries. Jev launched on September 15, and LangChain described it as having an “outsized response.” Anthropic’s Claude Plays Pokémon stream, running Opus 4.5, had not finished Red as of January, but with Jev playing and Claude writing rules, the developer achieved victory.
Model page: Jev →
Developer says Jev decision model beat Pokémon Red in under a week
Tom's Hardware · 27 September 2026
Loading the full article…
This text was published by Tom's Hardware and written by Shane Downing. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- Good Architecture deletes signals your agent depends on · 1 src
- Creator shares GPT-6 Astra review skill based on Arc · 1 src
- Business Insider staffers test Meta Muse and Instinct AI agents · 1 src
- Anthropic employees say they'd give AI agents about 30% of their book budget · 1 src
- OpenAI agents likely performed 16,500+ scans of UNCTADstat API · 1 src
Comments
via GitHub Discussions