CALM improves safety in text-to-image generation
The arXiv preprint titled 'Keep It CALM: Analyzing the Limits of Global Unsafety in Text‑to‑Image Generation' introduces a new training‑free safeguard for text‑to‑image models. The authors argue that existing global unsafe‑signal methods trade coverage for selectivity, failing to cover heterogeneous unsafe semantics while distorting benign prompts. CALM, or Counterfactual Adaptive Local…
Key points
- CALM is a training‑free safeguard that uses prompt‑local counterfactual correction instead of global unsafe‑signal removal.
- It matches unsafe‑benign anchors, edits token representations toward the safe side, and suppresses unsafe residual components.
- Evaluation shows CALM improves unsafe content suppression while preserving benign utility.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Google VP Yossi Matias says AI’s biggest impact may come from intersecting fields · 1 src
- Researchers adapt speech language model for simultaneous translation using prefix supervision · 1 src
- Study finds inductive prompting most consistent for LLM generalization in temporal extraction · 1 src
- AraBERT-based framework reaches 96.88% accuracy on Arabic DP ambiguity · 1 src
- Researchers introduce APDMem hierarchical memory for long-context LLM assistants · 1 src
Comments
via GitHub Discussions