DigestAI news desk

AI news, digested. Every story with its sources, every 30 minutes.

Researchupdated

Web-search decisions vary across ChatGPT, Claude, Grok, and DeepSeek agents, study says

A new arXiv paper investigates how conversational large‑language‑model agents use Web search. The authors examined four major platforms – ChatGPT, Claude, Grok and DeepSeek – by combining real‑world user interactions (in‑vivo) with controlled experiments using the same models via their APIs (in‑vitro). The study looks at when agents decide to invoke search, how they craft queries, the domains…

1 source primary source

Key points

  • Study examined Web‑search behavior of ChatGPT, Claude, Grok, and DeepSeek agents using real interactions and API experiments.
  • Search invocation frequency varied across platforms and did not consistently improve response quality.
  • Agents showed platform‑specific query strategies and domain preferences, with some uncited claims in responses.

The authors report that the frequency of Web‑search calls differs substantially across platforms and that more frequent searches do not automatically lead to higher response quality. Agents employ varied, sometimes complex, querying strategies, and each platform’s search engine tends to return results from preferred domains. While most responses are grounded in retrieved results, some claims rely on uncited sources, raising concerns about attribution and reliability. The findings suggest design implications for future AI agents and Web‑search tools optimized for conversational retrieval.

Read the original atarXiv cs.AI · by Mahsa Amani, Seungeon Lee, Abhisek Dash, Asmaa El Fraihi, Yunah Jang, Elisabeth Kirsten, Qinyuan Wu, Krishna P. Gummadi, Manish Gupta, Abhilasha Ravichander, Muhammad Bilal Zafar, Soumi Das primary sourceOpen source ↗
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories