SIGDIAL 2026

TRACER: Early Failure Detection for Task-Oriented Dialogue

Erfan Nourbakhsh1, Rocky Slavin1, Ke Yang1, Anthony Rios1

¹University of Texas at San Antonio (UTSA)

Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL 2026)
Atlanta, Georgia, USA

Introduction

Task Overview

Dialogues often fail before the final breakdown is obvious. TRACER forecasts eventual failure from partial context so recovery can fire earlier.

Task overview showing a conversation that breaks down.

Figure 1. Task overview. The example shows a conversation that breaks down.

Paper

Abstract

Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present TRACER, a method for early failure detection in task-oriented dialogue. TRACER predicts from a partial dialogue whether the full conversation will eventually fail by combining simple trajectory signals from belief-state changes with text representations of the evolving dialogue state. We evaluate the method in both oracle and generated belief-state settings, and test how well it works when only 25%, 50%, 75%, or 100% of the dialogue is visible. Across these settings, TRACER detects useful failure signals well before the end of the conversation and outperforms heuristic, classical, and single-stream baselines. Code: github.com/erfan-nourbakhsh/TRACER.

Method

TRACER Architecture

Two streams: RoBERTa over serialized belief states, plus turn-level trajectory dynamics. Fuse representations to score failure risk and trigger recovery.

Overview of TRACER with belief-state text stream and temporal trajectory stream.

Figure 2. Overview of TRACER. Partial dialogues are encoded by a belief-state text stream and a temporal trajectory stream. The text stream uses RoBERTa to encode serialized belief states, while the trajectory stream models turn-level dialogue-state dynamics. Their representations are fused to predict failure risk and decide whether to trigger recovery.

Results

MultiWOZ Failure Forecasting

TRACER leads on AUC-ROC and F1. Some heuristics get higher EDS only by triggering aggressively.

Category Model Belief AUC-ROC ↑ F1 ↑ EDS ↑ Det. Turn ↓
HeuristicNo interventionoracle.500.000.000∞
HeuristicFixed-schedule (every 3)oracle.499.236.6902.00
HeuristicFeature-threshold ensembleoracle.535.286.8640.25
Classicallogreg (engineered)oracle.600.351.7281.03
Zero-shotQwen2.5-7Boracle.531.321.7721.32
Few-shotLlama-3.1-8Boracle.492.337.9930.06
OursTRACERgenerated.617.370.7511.22
OursTRACERoracle.741†.463†.7271.36

Table 1. Main failure-forecasting results on MultiWOZ (selected rows). † Significantly better than best classical and best LLM baseline (p < 0.0001, paired bootstrap).

Ablation

Both Streams Matter

Fusion beats features-only and text-only under both generated and oracle belief states.

Model Belief AUC-ROC ↑ F1 ↑
Stream A only (features)generated.554.337
Stream A only (features)oracle.623.363
Stream B only (text)generated.583.332
Stream B only (text)oracle.692.432
TRACERgenerated.617.370
TRACERoracle.741.463

Table 2. Ablation on TRACER streams under oracle and generated belief-state inputs.

Threshold tradeoff between early detection and false positives.

Figure 3. Threshold tradeoff on the development set. Lower thresholds improve early detection but sharply increase false positives.

Partial Context

Signal Before the End

Fixed-context forecasting: useful signal already at 50% and 75% of the dialogue, not only at completion.

Fixed-context forecasting performance at 25%, 50%, 75%, and 100% of dialogue.

Figure 4. Fixed-context forecasting: TRACER recovers strong predictive signal from partial context, with substantial gains at 50% and 75% of the dialogue.

Conclusion

Main Takeaways

BibTeX

@inproceedings{nourbakhsh-etal-2026-tracer,
    title = "{TRACER}: Early Failure Detection for Task-Oriented Dialogue",
    author = "Nourbakhsh, Erfan  and
      Slavin, Rocky  and
      Yang, Ke  and
      Rios, Anthony",
    booktitle = "Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue",
    month = aug,
    year = "2026",
    address = "Atlanta, Georgia, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.sigdial-1.44/",
    pages = "614--635",
}