- Introduction
- News
- How It Works
- Installation
- Inference
- Evaluation
- Centroid Cache
- Output Structure
- Citation
- Acknowledgement
Entropy Centroid is a training-free test-time compute method that selects the best trajectory from N sampled responses using entropy dynamics.
Key features:
- Runs on top of any vLLM-compatible model — no fine-tuning required
- Supports math, logic, code, and agent benchmarks
- Unified evaluation pipeline for all dataset types
- [2026/04] Initial release of inference and evaluation pipeline.
- Inference — sample N trajectories per problem, recording per-token top 10 candidates in the vocabulary with the highest probability
- Centroid cache — compute a scalar centroid score per trajectory from High Entropy Phase (HEP)
- Evaluation — pick the trajectory with the lowest centroid score as the final answer and score the selected answer against ground truth.
pip install -r requirements.txtFor tau2-bench (agent inference), install the submodule separately:
cd tau2-bench && pip install -e .Config files live in scripts/configs/. Set DATASET_TYPE at the top of run_batch.sh, then run:
bash scripts/run_batch.shSupported dataset types:
DATASET_TYPE |
Datasets | Config |
|---|---|---|
aime |
AIME 2025 | config_aime_2025yaml |
minerva |
Minerva | config_minerva.yaml |
livecodebench |
LiveCodeBench | config_livecodebench.yaml |
bigcodebench |
BigCodeBench | config_bigcodebench.yaml |
synlogic |
SynLogic | config_synlogic.yaml |
Key parameters in run_batch.sh:
DATASET_TYPE="math"
INFERENCE_MODELS=("Qwen/QwQ-32B")
GPU_IDS="0,1,2,3,4,5,6,7"
TENSOR_PARALLEL_SIZE=8
TRAJECTORIES_PER_SAMPLE=64 # N in Best-of-N
TEMPERATURE=0.7
MAX_TOKENS=32768tau2-bench uses a separate entry point (tau2.cli) with a local vLLM server:
bash scripts/run_tau2_bench.sh
# Options:
bash scripts/run_tau2_bench.sh --model-idx 0 # single model by index
bash scripts/run_tau2_bench.sh --domain retail # single domain
bash scripts/run_tau2_bench.sh --skip-existing # resume interrupted runs
bash scripts/run_tau2_bench.sh --dry-run # 2 tasks per domain onlyOutput: outputs/tau2_bench/<model>/{airline,retail,telecom}/entropy_results.jsonl
bash scripts/run_greedy_baseline.sh
# Scope to specific benchmark groups:
BENCHMARKS="math" bash scripts/run_greedy_baseline.sh
BENCHMARKS="logic" bash scripts/run_greedy_baseline.sh
BENCHMARKS="code" bash scripts/run_greedy_baseline.sh
# Evaluate existing results without re-running inference:
EVAL_ONLY=true bash scripts/run_greedy_baseline.sh# Lowest-centroid selection (main method)
python evaluate_answers.py --result_dir <RESULT_DIR> --selection lowest_centroid
# Batch over multiple directories
python evaluate_unified.py --batch --pattern "outputs/results/config_aime*"
# Other selection methods
python evaluate_answers.py --result_dir <RESULT_DIR> --selection majority_voting
python evaluate_answers.py --result_dir <RESULT_DIR> --selection pass_at_k --pass_k 1,5,10,32python evaluate_livecodebench.py --result_dir <RESULT_DIR>
python evaluate_livecodebench.py --batch --pattern "outputs/results/*livecodebench*"python evaluate_bigcodebench.py --result_dir <RESULT_DIR>
python evaluate_bigcodebench.py --batch --pattern "outputs/results/bigcodebench_*"python evaluate_tau2_bench.py --result_dir <RESULT_DIR>
python evaluate_tau2_bench.py --batch --batch_base_dir outputs/tau2_benchAll evaluators accept the following centroid parameters:
--centroid_top_percent 1.0 # fraction of top-entropy tokens used as step boundaries
--centroid_bottom_percent 80.0 # fraction of low-entropy tokens used as centroid anchors
--centroid_consecutive_low 2 # consecutive low-entropy token threshold
--centroid_method hep # hep | raw_entropy
--force # recompute existing cachesThe centroid cache (trajectory_centroid_cache_*.json) is built automatically during evaluation. To pre-build it in parallel across a parameter sweep (useful before running large-scale evaluations):
# Default sweep: top=[1,3,5] × bottom=[30,50,80] × cons=[2,3,5] = 27 combos per dir
bash scripts/build_centroid_cache.sh
# Scope to specific directories
bash scripts/build_centroid_cache.sh --dir outputs/results/config_aime_2025_QwQ-32B_*/
# Extended sweep
bash scripts/build_centroid_cache.sh --extended
# Options
bash scripts/build_centroid_cache.sh --raw # raw_entropy method only
bash scripts/build_centroid_cache.sh --force # recompute existing caches
bash scripts/build_centroid_cache.sh --max_jobs 32 # concurrency limit
bash scripts/build_centroid_cache.sh --filter_llm_error # tau2: skip llm_error trajectoriestau2-bench directories (containing entropy_results.jsonl) are detected automatically and routed through the appropriate preparation pipeline.
outputs/
├── results/ # math / logic / code inference
│ └── <dataset>_<model>_<ts>/
│ ├── entropy_results.json
│ ├── evaluation_cache.json
│ ├── trajectory_centroid_cache_*.json
│ └── eval_lowest_centroid/
│ ├── answer_evaluation_summary.json
│ └── trajectory_selections.json
├── greedy/ # greedy baseline
│ └── greedy_<dataset>_<model>_<ts>/
├── tau2_bench/ # tau2-bench agent inference
│ └── <model>/
│ └── {airline,retail,telecom}/
│ └── entropy_results.jsonl
└── logs/
If you find this work useful, please cite:
@article{zhao2026entropycentroids,
title={Entropy Centroids as Intrinsic Rewards for Test-Time Scaling},
author={Wenshuo Zhao and Qi Zhu and Xingshan Zeng and Fei Mi and Lifeng Shang and Yiren Feng},
journal={arXiv preprint arXiv:2604.26173},
year={2026}
}