Researchers introduce JEVal benchmark to test General decision models
A new paper on arXiv compares general decision models like Jev with traditional LLMs. The authors created JEVal, a benchmark of 11,257 instances across 36 datasets and 10 domains, to measure performance in structured judgment and selection tasks.
Key points
- JEVal benchmark tests 25 model configurations across 10 application domains with 11,257 instances
- General decision models outperform LLMs in evidence-based decisions but overestimate certainty and fail in specialist tasks
- InnerJev-27B matches Jev’s performance on JEVal while answering queries in 0.1 seconds
The study finds that decision models excel when decisions rely on clear evidence but struggle with specialist knowledge or uncertainty estimation. In dynamic systems, faster local decisions reduce processing time but increase errors over long trajectories. On social simulations, these models match LLMs in individual predictions but lag in user profiling and exhibit bias. The paper also introduces InnerJev-4B and InnerJev-27B, optimized models that distill reasoning into a single-pass decision, achieving Jev-level performance on JEVal in about 0.1 seconds.
Model page: Jev →
The story so far
2 episodes →- Researchers introduce JEVal benchmark to test General decision modelsthis story
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic study: task understanding beats job title for AI success · 1 src
- Researchers test fixed token codes for language models at 100B-token scale · 1 src
- Baibaichuchu at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text? · 1 src
- Researchers release OncoNoteBERT for oncology note processing · 1 src
- Study finds brain-alignment and Cross-Lingual scores can mislead when probes fail · 1 src
Comments
via GitHub Discussions