AACL-IJCNLP 2026 Main

Prompting Underestimates LLM Capability for Time Series Classification

Dan Schumacher1, Erfan Nourbakhsh1, Rocky Slavin1, Anthony Rios1

ยนUniversity of Texas at San Antonio (UTSA)

The 5th Asia-Pacific Chapter of the Association for Computational Linguistics
& the 15th International Joint Conference on Natural Language Processing (AACL-IJCNLP 2026)
Main Conference Paper

Introduction

Main Idea

Prompt pipelines look weak on time-series classification. Same model internals, read with a linear probe, tell a different story.

Main idea: prompting underestimates what probes recover from LLM representations.

Figure 1. Main idea: prompt-based generation underestimates representational capacity revealed by linear probing.

Paper

Abstract

Prompt-based prediction pipelines suggest that large language models (LLMs) perform poorly on time-series classification, raising doubts about whether they encode meaningful temporal structure. We show that this conclusion reflects limitations of prompt-based generation rather than the model's representational capacity by directly comparing prompt outputs with linear probes over the same internal representations while controlling for important confounding factors. While prompting performs near chance, linear probes improve average F1 from 0.15-0.26 to 0.61-0.67, often matching or exceeding specialized state-of-the-art (SoTA) time-series models. Layer-wise analyses further show that class-discriminative time-series information emerges in early transformer layers and is amplified with multimodal inputs and the presence of pretrained weights. Together, these results demonstrate a systematic mismatch between what LLMs internally represent and what prompt-based prediction pipelines reveal.

Setup

Prompting vs Probing

Digits (d), visualizations (v), or both (dv). Same prompt used for generation (PBC) and representation extraction (RP). No fine-tuning, no extra TS embedders.

Overview of prompting versus probing paradigms for time series with digit and visual modalities.

Figure 2. Overview of prompting vs. probing paradigms. Raw time series are transformed into digit-based text representations and visualizations which are incorporated into prompts for direct prediction or probing.

Results

Probe Beats Prompt

On dv, probes jump F1 from ~0.15-0.26 to ~0.61-0.67 and compete with Attend / TS2Vec / Moment. Prompting stays near chance.

Method CTUEMGHADHARRWCTEEAvg
Attend.848.985.732.933.726.524.791
TS2Vec.631.933.667.865.672.794.761
Moment.6601.00.671.878.778.564.758
OneFitsAll.587.296.681.885.790.475.615
InstructTime.627.167.278.527.434.235.378
Llama RP.6401.00.406.723.672.223.611
Llama PBC.356.175.033.111.190.036.150
Mistral RP.684.933.438.730.645.462.649
Mistral PBC.548.393.021.127.190.288.261
Qwen RP.6761.00.374.737.656.592.672
Qwen PBC.324.413.016.174.222.304.242

Table 1. Macro F1 across datasets. Prompt (PBC) and Probe (RP) results for the dv modality. Bold: probe averages.

Analysis

Where Signal Emerges

Class-discriminative structure appears early. Multimodal inputs and pretrained weights amplify deeper layers.

Layer-wise probe macro F1 for Mistral and Qwen across modalities.

Figure 3. Layer-wise probe macro F1 for Mistral and Qwen across modalities.

Layer-wise F1 for Qwen versus randomly initialized weights.

Figure 4. Layer-wise macro F1 of probes for Qwen versus the same model with randomly initialized weights.

Reliability

Prompting Is Unstable

Prompt wording and sampling swing PBC scores. Visual-only prompting is a bit more stable, but still far below probes.

Prompt-based macro F1 distributions by input modality.

Figure 5. Prompt-based macro F1 distributions by input modalities, showing a bit more stable performance for visual-only prompting.

Conclusion

Main Takeaways

BibTeX

@misc{schumacher2026promptingunderestimatesllmcapability,
      title={Prompting Underestimates LLM Capability for Time Series Classification},
      author={Dan Schumacher and Erfan Nourbakhsh and Rocky Slavin and Anthony Rios},
      year={2026},
      eprint={2601.03464},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2601.03464},
}