ECCV 2026 arXiv v1 · 28 Aug 2026

ClearText-Video

A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

A quality-controlled benchmark that connects video restoration with the harder question: can multimodal models still read and reason about the same text when video quality changes?

01 Benchmark overview
Overview of the ClearText-Video benchmark linking restoration evaluation, high-quality video QA, degraded-quality video QA, and restored-quality video QA.
One content instance, matched across high-quality, degraded and restored video evidence.
4,639text-rich videos
550K+video frames
1.6Mverified annotations
220K+spatial & temporal QA

Why ClearText-Video

Video quality changes the evidence—not only the pixels.

Real-world videos contain motion blur, compression, noise and low-resolution text. ClearText-Video makes those changes measurable by pairing every high-quality source with content-matched degraded and restored variants.

It unifies two task families—Text-Centric Video Restoration and Multi-Quality VideoQA—so perceptual enhancement can be evaluated alongside text fidelity and downstream reasoning.

Does a video that looks better also preserve the textual evidence a multimodal model needs?

Dataset & tasks

Built for text in motion,
not text in isolation.

Chinese and English scene text captured in real indoor and outdoor environments, with aligned quality variants and two rounds of expert verification.

1

Source videos

Real-world, text-rich egocentric clips from offline and online sources.

2

Human verification

Text boxes, transcripts, captions and QA reviewed in two rounds.

3

Matched quality

HQ originals paired with controlled degradations and restored outputs.

4

VideoQA

Spatial and temporal questions test reading, grounding and reasoning.

Content-matched quality regimes
HQ
High qualityOriginal evidence
DQ
Low resolutionBicubic ×4
DQ
BlurLocally variant blur
RQ
DOVERestored evidence
RQ
MIMORestored evidence
RQ
S3DIFFRestored evidence

Public release scope: the Hugging Face release currently provides GT, blur and downsample_x4. Full-paper RQ variants are evaluation conditions reported in the paper.

02 Bilingual real-world scenes
Twelve English and Chinese scene-text examples from trucks, signs, transport, airports and storefronts.
Text appears on vehicles, directional signs, storefronts, billboards and other everyday carriers.
03 Two-round annotation
Collection, labeling, mixed-expert QA processing, inspection and human correction pipeline.
Training 4,327 videos

Large-scale supervision for restoration-aware and quality-robust model development.

Testing 312 videos

74% offline / 26% online, with balanced Chinese and English coverage by source.

Difficulty 46 / 37 / 17%

Easy, medium and hard questions for controlled difficulty analysis.

04 Captions, text and scene coverage
Word clouds and distributions of caption language, scene text, scene categories and text carrier types.

Benchmark

Restoration improves pixels—
not always the evidence.

ClearText-Video evaluates visual quality, text fidelity and multimodal reasoning together instead of treating restoration as an endpoint.

18restoration methods
16multimodal LLMs
6spatial quality regimes
4spatial metrics
05 Accuracy across quality regimes
Radar charts comparing accuracy and unbiased accuracy of sixteen multimodal models across HQ, degraded and restored quality conditions.
Metrics: Accuracy (Acc), Unbiased Accuracy (UAcc), Overconfidence (OC) and Answer Abstention (Abs).
01

Blur is the harder degradation

Across 16 MLLMs, spatial accuracy drops 6.05 points under blur versus 3.14 points under low resolution.

02

Best spatial HQ · 71.67%

Gemini-2.5-pro leads high-quality spatial VideoQA and also leads DQ-Low_res and RQ-S3DIFF.

03

Best temporal HQ · 60.02%

Claude-Sonnet-4.6 ranks first across all five temporal quality conditions.

04

Fine-tuning helps—but shifts reliability

Qwen2.5-VL-7B-SFT gains 6.31–10.11 points over its base model, while accuracy and uncertainty-aware reliability do not always move together.

Spatial VideoQA

Best accuracy by condition

Final paper results (%)
ConditionBest modelAcc
HQGemini-2.5-pro71.67
DQ-Low_resGemini-2.5-pro65.00
DQ-BlurClaude-Sonnet-4.660.00
RQ-DOVEGemini-2.5-flash66.67
RQ-MIMOClaude-Sonnet-4.670.00
RQ-S3DIFFGemini-2.5-pro70.00
Temporal VideoQA

Claude-Sonnet-4.6

Best across all conditions (%)
ConditionAccuracy
HQ60.02
DQ-Blur59.45
RQ-MIMO59.34
RQ-S3DIFF58.88
RQ-DOVE60.37

Failure modes

Restored outputs can alter text.

Perceptual enhancement can introduce residual character errors or hallucinations that change the evidence used for downstream answers.

Examples of text hallucination, character distortion, localization bias, reasoning errors and temporal aggregation errors after restoration.

Qualitative restoration

Read the details that global image metrics miss.

ZH Chinese scene text
Chinese scene-text restoration comparison across high-quality, low-resolution and restored outputs.
EN English scene text
English scene-text restoration comparison across high-quality, low-resolution and restored outputs.

Resources

Start with the benchmark.

Paper, public data and reproducible evaluation code are collected here.

Citation

If this benchmark supports your research, please cite:

@inproceedings{li2026cleartextvideo,
  title     = {ClearText-Video: A Large-Scale Text-Centric Video Dataset
               Bridging Video Restoration and Scene-Text Enhancement},
  author    = {Li, Jinlong and Ding, Jiaming and Lu, Dingfu and Hsiu, Malcolm
               and Ke, Chuang and Yang, Kangning and Guan, Bochen and Fu, Lan
               and Cai, Jie and Sun, Huiming and Meng, Zibo},
  booktitle = {European Conference on Computer Vision},
  eprint    = {2608.28784},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url       = {https://arxiv.org/abs/2608.28784},
  year      = {2026}
}

Authors

Jinlong Li*†, Jiaming Ding*, Dingfu Lu, Malcolm Hsiu, Chuang Ke, Kangning Yang, Bochen Guan, Lan Fu, Jie Cai, Huiming Sun, Zibo Meng

OPPO US AI Center · University of Wisconsin–Madison · University of California San Diego

* Equal contribution · Corresponding author · Work done during internships at OPPO US AI Center