DART: A Training-Free Router for Adaptive Thinking Budgets in Hybrid Reasoning Models
ORIGINAL / DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models
This paper presents DART, a training-free routing framework for hybrid reasoning models, which determines whether to answer directly or allocate more thinking budget based on the agreement of two cheap no-think drafts. It adapts computation to problem difficulty without labeled data, improving efficiency while preserving accuracy.
01 ABSTRACT
The paper introduces DART, which uses consensus and entropy from two no-think drafts to route queries. The authors report that DART maintains or improves accuracy while reducing thinking tokens by 32-73%, with up to +9.0 points on olympiad-level math and +22.5 on code. The signal generalizes across model scales and settings. These claims are based on the paper's experiments; details are limited.
02 KEY FINDINGS
- DART is training-free, requiring no labels or gradient updates.
- It decides based on consensus of two no-think drafts.
- It reduces thinking tokens by 32-73% while maintaining or improving accuracy.
- Accuracy gains up to 9.0 on math and 22.5 on code.
- Signal works across model sizes (0.6B-32B) and API settings.
AI GENERATED SUMMARY / DISCOVERED BY ARXIV CS.CL