DigestAI news desk
OpenAI board member warns company is not on track to prevent catastrophic AI loss of control OpenAI Unveils GPT‑6 Astra: Record‑Breaking 3D Rendering, Loop‑Transformer Architecture OpenAI launches Agents API beta for long-running cloud agents OpenAI solves Navier-Stokes problem, sparking academic controversy over data use OpenAI Introduces ChatGPT for Financial Services The Waymo effect: AI making research less collaborative RTK Token Savings Debunked: Cost Benchmarks Disagree Ypsilanti Township Residents Protest Nuclear AI Data Center Proposal
Research updated 2 min read

DiscoSign introduces discourse-aware translation from text to ASL gloss using LLMs

DiscoSign is a new framework that extends text‑to‑sign‑language gloss translation beyond the sentence level by incorporating discourse‑level cues. Built on a modular large language model pipeline, it tackles three linguistic challenges: spatial coreference (keeping entity locations consistent across sentences), Question‑Answer Clauses that serve specific discourse functions, and concept‑gloss…

1 source primary source

Key points

  • DiscoSign adds spatial coreference, QAC handling, and concept‑gloss consistency to sign‑language gloss translation
  • Experiments show improved entity tracking and spatial consistency over sentence‑only baselines while keeping single‑sentence quality
  • New evaluation metrics assess discourse coherence, filling a gap in sign‑language translation assessment

The authors evaluate DiscoSign on both sentence‑level and discourse‑level datasets, reporting marked gains in spatial consistency and entity tracking compared with traditional sentence‑only systems, while preserving competitive single‑sentence gloss quality. To measure these improvements, they also introduce a suite of novel metrics that capture discourse coherence, addressing a long‑standing gap in sign‑language translation evaluation.

By providing the first systematic approach and evaluation suite for discourse‑aware sign‑language gloss translation, DiscoSign paves the way for more natural, context‑sensitive AI‑driven interpretation tools for the Deaf and Hard‑of‑Hearing community.

Full story from Apple Machine Learning Research primary source Open source ↗

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Apple Machine Learning Research · 11 September 2026

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

AuthorsVasileios Baltatzis‡, Mert Inan‡†, Connor Gillis, Raja Kushalnagar§, Lorna Quandt§**, Leah Findlater, Colin Lea

Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving specific discourse functions; and (iii) concept-gloss consistency, ensuring stable mappings between English concepts and American Sign Language (ASL) signs. Traditional translation metrics fail to capture discourse-level quality, so we introduce a suite of novel evaluation metrics designed to assess each dimension of discourse coherence addressed by our framework. Experiments on sentence-level and discourse-level datasets show that our approach for discourse-aware processing significantly improves spatial consistency and entity tracking relative to sentence-only translation, while maintaining competitive single-sentence gloss translation quality. Our work establishes the first systematic framework for discourse-level text to sign language gloss translation with corresponding evaluation methodology.

Bootstrapping Sign Language Annotations with Sign Language Models

April 30, 2026research area Accessibility, research area Computer Visionconference CVPR

AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of data but remain only partially annotated and thus underutilized, in part due to the prohibitive costs of annotating at this scale. In this work, we develop a pseudo-annotation pipeline that takes signed video and English as input and outputs a ranked…

Towards AI-Driven Sign Language Generation with Non-Manual Markers

March 7, 2025research area Accessibility, research area Human-Computer Interactionconference CHI

Sign languages are essential for the Deaf and Hard-of-Hearing (DHH) community. Sign language generation systems have the potential to support communication by translating from written languages, such as English, into signed videos. However, current systems often fail to meet user needs due to poor translation of grammatical structures, the absence of facial cues and body language, and insufficient visual and motion fidelity. We address these…

This text was published by Apple Machine Learning Research . It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

Related stories