Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

Paper

centroid-intro-figure_00

Table of Contents

Introduction

Entropy Centroid is a training-free test-time compute method that selects the best trajectory from N sampled responses using entropy dynamics.

Key features:

  • Runs on top of any vLLM-compatible model — no fine-tuning required
  • Supports math, logic, code, and agent benchmarks
  • Unified evaluation pipeline for all dataset types

News

  • [2026/04] Initial release of inference and evaluation pipeline.

How It Works

  1. Inference — sample N trajectories per problem, recording per-token top 10 candidates in the vocabulary with the highest probability
  2. Centroid cache — compute a scalar centroid score per trajectory from High Entropy Phase (HEP)
  3. Evaluation — pick the trajectory with the lowest centroid score as the final answer and score the selected answer against ground truth.

Installation

pip install -r requirements.txt

For tau2-bench (agent inference), install the submodule separately:

cd tau2-bench && pip install -e .

Inference

Math / Logic / Code

Config files live in scripts/configs/. Set DATASET_TYPE at the top of run_batch.sh, then run:

bash scripts/run_batch.sh

Supported dataset types:

DATASET_TYPE Datasets Config
aime AIME 2025 config_aime_2025yaml
minerva Minerva config_minerva.yaml
livecodebench LiveCodeBench config_livecodebench.yaml
bigcodebench BigCodeBench config_bigcodebench.yaml
synlogic SynLogic config_synlogic.yaml

Key parameters in run_batch.sh:

DATASET_TYPE="math"
INFERENCE_MODELS=("Qwen/QwQ-32B")
GPU_IDS="0,1,2,3,4,5,6,7"
TENSOR_PARALLEL_SIZE=8
TRAJECTORIES_PER_SAMPLE=64   # N in Best-of-N
TEMPERATURE=0.7
MAX_TOKENS=32768

tau2-bench (Agent Inference)

tau2-bench uses a separate entry point (tau2.cli) with a local vLLM server:

bash scripts/run_tau2_bench.sh

# Options:
bash scripts/run_tau2_bench.sh --model-idx 0      # single model by index
bash scripts/run_tau2_bench.sh --domain retail    # single domain
bash scripts/run_tau2_bench.sh --skip-existing    # resume interrupted runs
bash scripts/run_tau2_bench.sh --dry-run          # 2 tasks per domain only

Output: outputs/tau2_bench/<model>/{airline,retail,telecom}/entropy_results.jsonl

Greedy Baseline

bash scripts/run_greedy_baseline.sh

# Scope to specific benchmark groups:
BENCHMARKS="math"  bash scripts/run_greedy_baseline.sh
BENCHMARKS="logic" bash scripts/run_greedy_baseline.sh
BENCHMARKS="code"  bash scripts/run_greedy_baseline.sh

# Evaluate existing results without re-running inference:
EVAL_ONLY=true bash scripts/run_greedy_baseline.sh

Evaluation

Math / Logic

# Lowest-centroid selection (main method)
python evaluate_answers.py --result_dir <RESULT_DIR> --selection lowest_centroid

# Batch over multiple directories
python evaluate_unified.py --batch --pattern "outputs/results/config_aime*"

# Other selection methods
python evaluate_answers.py --result_dir <RESULT_DIR> --selection majority_voting
python evaluate_answers.py --result_dir <RESULT_DIR> --selection pass_at_k --pass_k 1,5,10,32

LiveCodeBench

python evaluate_livecodebench.py --result_dir <RESULT_DIR>
python evaluate_livecodebench.py --batch --pattern "outputs/results/*livecodebench*"

BigCodeBench

python evaluate_bigcodebench.py --result_dir <RESULT_DIR>
python evaluate_bigcodebench.py --batch --pattern "outputs/results/bigcodebench_*"

tau2-bench

python evaluate_tau2_bench.py --result_dir <RESULT_DIR>
python evaluate_tau2_bench.py --batch --batch_base_dir outputs/tau2_bench

Centroid Parameters

All evaluators accept the following centroid parameters:

--centroid_top_percent 1.0       # fraction of top-entropy tokens used as step boundaries
--centroid_bottom_percent 80.0   # fraction of low-entropy tokens used as centroid anchors
--centroid_consecutive_low 2     # consecutive low-entropy token threshold
--centroid_method hep            # hep | raw_entropy
--force                          # recompute existing caches

Centroid Cache

The centroid cache (trajectory_centroid_cache_*.json) is built automatically during evaluation. To pre-build it in parallel across a parameter sweep (useful before running large-scale evaluations):

# Default sweep: top=[1,3,5] × bottom=[30,50,80] × cons=[2,3,5] = 27 combos per dir
bash scripts/build_centroid_cache.sh

# Scope to specific directories
bash scripts/build_centroid_cache.sh --dir outputs/results/config_aime_2025_QwQ-32B_*/

# Extended sweep
bash scripts/build_centroid_cache.sh --extended

# Options
bash scripts/build_centroid_cache.sh --raw                # raw_entropy method only
bash scripts/build_centroid_cache.sh --force              # recompute existing caches
bash scripts/build_centroid_cache.sh --max_jobs 32        # concurrency limit
bash scripts/build_centroid_cache.sh --filter_llm_error   # tau2: skip llm_error trajectories

tau2-bench directories (containing entropy_results.jsonl) are detected automatically and routed through the appropriate preparation pipeline.

Output Structure

outputs/
├── results/                          # math / logic / code inference
│   └── <dataset>_<model>_<ts>/
│       ├── entropy_results.json
│       ├── evaluation_cache.json
│       ├── trajectory_centroid_cache_*.json
│       └── eval_lowest_centroid/
│           ├── answer_evaluation_summary.json
│           └── trajectory_selections.json
├── greedy/                           # greedy baseline
│   └── greedy_<dataset>_<model>_<ts>/
├── tau2_bench/                       # tau2-bench agent inference
│   └── <model>/
│       └── {airline,retail,telecom}/
│           └── entropy_results.jsonl
└── logs/

Citation

If you find this work useful, please cite:

@article{zhao2026entropycentroids,
      title={Entropy Centroids as Intrinsic Rewards for Test-Time Scaling}, 
      author={Wenshuo Zhao and Qi Zhu and Xingshan Zeng and Fei Mi and Lifeng Shang and Yiren Feng},
      journal={arXiv preprint arXiv:2604.26173},
      year={2026}
}

About

No description, website, or topics provided.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages