A lightweight benchmark for evaluating code-oriented language models on task performance, latency, and token efficiency.
This project compares multiple open models on a code agent style evaluation and visualizes the tradeoffs between quality and speed.
The goal of this benchmark is simple:
- Measure functional correctness using Pass@1
- Track average latency per task
- Record token usage
- Make tradeoffs visible
The current comparison includes:
meta-llama/llama-4-scout-17b-16e-instructllama-3.1-8b-instantallam-2-7bopenai/gpt-oss-20b
The results are summarized both in CSV format and in a visual performance vs latency chart.
Files:
-
eval_benchmark.ipynb
Main evaluation notebook. Runs model calls, captures latency and tokens, and computes summary metrics. -
detailed_results.csv
Row-level results for each task and model. -
metrics_summary.csv
Aggregated metrics per model. -
performance_vs_latency.png
Bubble chart showing quality vs speed tradeoffs.
Percentage of tasks solved correctly on the first attempt.
Higher means better reliability without retries.
Mean time per task.
Lower is better for interactive or agent-based systems.
Average number of tokens consumed per task.
Useful for estimating cost and efficiency.
| Model | Pass@1 | Avg Latency (s) | Avg Tokens |
|---|---|---|---|
| meta-llama/llama-4-scout-17b-16e-instruct | 100% | ~0.48 | ~305 |
| llama-3.1-8b-instant | 50% | ~0.13 | ~215 |
| allam-2-7b | 10% | ~0.34 | ~280 |
| openai/gpt-oss-20b | 10% | ~0.41 | ~490 |
| Benchmark | Metric | Score |
|---|---|---|
| GSM8K (n=50) | Accuracy | 0.540 |
| HumanEval-style (n=10) | pass@1 | 0.700 |
| LLM-as-Judge (n=10) | Avg score | 8.2/10 |
| ↳ Math reasoning | score | 8.5/10 |
| ↳ Code generation | score | 8.0/10 |
| ↳ Logical reasoning | score | 9.0/10 |
| ↳ Instruction following | score | 8.0/10 |
| ↳ Knowledge | score | 7.5/10 |
Full results: results/base_model/eval_results.json
llama-3.1-8b-instantis the fastest model in this benchmark.meta-llama/llama-4-scout-17b-16e-instructachieved the highest Pass@1 score.openai/gpt-oss-20bconsumed the most tokens in this run.- Smaller models tend to be faster, but accuracy varies significantly depending on task complexity.
The performance_vs_latency.png file plots:
- X-axis: Average latency
- Y-axis: Pass@1
- Bubble size: Average tokens
This makes it easy to compare speed, quality, and efficiency at a glance.
A ReAct-style harness that evaluates agents on 14 tasks requiring multi-turn tool calls. The agent thinks, calls a tool, receives an [OBS] result, and iterates until it emits <final_answer>.
| Tool | Description |
|---|---|
python_exec |
Subprocess Python sandbox — arbitrary computation |
calculator |
Safe expression evaluator (no arbitrary code) |
lookup |
Retrieves constants from a knowledge base (π, speed of light, etc.) |
| Category | Tasks | What it tests |
|---|---|---|
computation |
5 | Single tool call for direct calculation |
multi_step |
5 | Chained reasoning: digit sum of factorial, Collatz, geometric series |
knowledge_compute |
2 | Lookup → calculator: circle area with exact π, light travel time |
error_recovery |
2 | First call fails (sqrt(-1), div-by-0) → agent retries correct call |
python eval_tool_use.py --mode demoMeasured results (deterministic ceiling, 14 tasks):
| Metric | Value |
|---|---|
| Task success | 100.0% (14/14) |
| Avg steps/task | 1.3 |
| Tool accuracy | 88.9% (expected errors counted) |
| Error recovery | 100.0% (2/2 error tasks recovered) |
| Avg latency | 0.015 s (subprocess only) |
Per-category: computation 100% · multi_step 100% · knowledge_compute 100% · error_recovery 100%
# Requires GPU + Qwen2.5-7B-Instruct (or any HF causal LM)
CUDA_VISIBLE_DEVICES=0 python eval_tool_use.py --mode hf --model Qwen/Qwen2.5-7B-InstructUses apply_chat_template for multi-turn ReAct context so instruct-tuned models receive proper <|im_start|> formatting.
Measured results (Qwen2.5-7B-Instruct, 14 tasks, A30):
| Metric | Value |
|---|---|
| Task success | 78.6% (11/14) |
| Avg steps/task | 1.7 |
| Tool accuracy | 91.7% |
| Error recovery | 50.0% (1/2) |
| Avg latency | 8.2 s/task |
Per-category: computation 100% · multi_step 40% · knowledge_compute 100% · error_recovery 100%
Results are written to results/tool_use_results.json.
Add a task by appending to TASKS:
Task("my_task", "Compute the 20th prime number.", "73", category="computation")Add a tool by registering in TOOLS and adding a description to TOOL_DESCRIPTIONS.
A new multi-axis harness (eval_benchmarks.py) extends the original latency/pass@1 benchmark with:
| Benchmark | Task | Metric | How |
|---|---|---|---|
| GSM8K | Math reasoning (7,473 problems) | Exact-match accuracy | Numeric extraction + ±1e-4 tolerance |
| HumanEval-style | Code generation (10 representative problems) | pass@1 | Subprocess execution + test assertion |
| LLM-as-Judge | Open-ended quality (10 questions) | Score 0–10 | Self-judge via structured prompt |
- List manipulation (
has_close_elements,intersperse,rolling_max) - String parsing (
separate_paren_groups,parse_nested_parens,filter_by_substring) - Numeric (
truncate_number,below_zero,mean_absolute_deviation,sum_product)
- Math reasoning, Code generation, Logical reasoning, Instruction following, Knowledge
pip install torch transformers peft datasets
# Evaluate base model
python eval_benchmarks.py \
--base_model Qwen/Qwen2.5-7B-Instruct \
--benchmarks gsm8k humaneval llm_judge \
--gsm8k_n 100 --humaneval_n 10 --judge_n 10 \
--output_dir results/base
# Evaluate fine-tuned model (with LoRA adapter)
python eval_benchmarks.py \
--base_model Qwen/Qwen2.5-7B-Instruct \
--adapter_path /path/to/lora/adapter \
--benchmarks gsm8k humaneval llm_judge \
--output_dir results/finetunedOutputs: results/eval_results.json + results/eval_report.md
pip install pandas matplotlib
# You may also need SDKs or API clients for the models you are evaluating.jupyter notebook eval_benchmark.ipynbThe notebook will load tasks, query each model, measure latency and token usage, compute Pass@1, and export detailed_results.csv, metrics_summary.csv, and performance_vs_latency.png.
Add its inference call in eval_benchmark.ipynb, capture model output, latency, and token usage, re-run, and regenerate the summary CSV and visualization. Keep logging consistent so comparisons remain fair.
Extends the single ReAct agent with a three-agent orchestration loop that addresses two failure modes of the single-agent setup: hitting the step limit on multi-step tasks, and producing wrong-but-confident answers with no verification.
Planner — decomposes the task into a concrete ordered plan (category-aware templates in demo mode)
Executor — runs the ReAct loop anchored to the plan; knows which subgoal it's working on
Critic — independently verifies the answer; on rejection, injects corrective feedback and
triggers one retry
Demo mode (no GPU, CI-safe):
python multi_agent_pipeline.py --mode demoCompare single-agent vs multi-agent (real GPU):
CUDA_VISIBLE_DEVICES=0 python multi_agent_pipeline.py --mode hf --model Qwen/Qwen2.5-7B-Instruct --compareReal model (eval only):
CUDA_VISIBLE_DEVICES=0 python multi_agent_pipeline.py --mode hf --model Qwen/Qwen2.5-7B-InstructThe pipeline reuses all tool infrastructure and tasks from eval_tool_use.py. The Critic uses numeric equality checking in demo mode and an LLM-based judge prompt (with apply_chat_template) in real mode.
| Agent | Task success | Retry rate | Critic approval | Avg steps |
|---|---|---|---|---|
| Single-agent (ReAct) | 78.6% (11/14) | — | — | 1.7 |
| Multi-agent (Planner→Executor→Critic) | 85.7% (12/14) | 64.3% | 35.7% | — |
| Delta | +7.1 pp |
The Critic flags incorrect answers on 9 of 14 tasks; the corrective retry rescues 1 additional task. Multi-step reasoning (multi_step category) is where the Planner's explicit decomposition provides the most benefit.
Wraps lm-evaluation-harness to run standard tasks (GSM8K, HumanEval, MMLU, HellaSwag) and compare them directly against the custom eval numbers in metrics_summary.csv.
pip install lm-eval# Run standard harness on GSM8K + HumanEval (500 examples each)
python lm_eval_integration.py \
--model Qwen/Qwen2.5-7B-Instruct \
--tasks gsm8k humaneval \
--limit 500
# Run and compare against custom eval CSV
python lm_eval_integration.py \
--model Qwen/Qwen2.5-7B-Instruct \
--tasks gsm8k humaneval mmlu \
--limit 500 \
--compare --custom_csv metrics_summary.csvTask lm-eval score Custom eval score Delta
GSM8K 0.538 0.540 -0.002
HumanEval 0.695 0.700 -0.005
MMLU 0.612 — —
The comparison table is also saved to results/comparison_table.csv.
Benchmark contamination occurs when eval questions (or near-duplicates) appear verbatim in the training corpus, artificially inflating scores. This script computes character-level n-gram overlap between eval questions and a reference corpus (UltraFeedback by default).
A question is flagged as contaminated if more than threshold (default 20%) of its 13-grams appear in the corpus — consistent with the methodology used in Gopher and Chinchilla.
- A contaminated benchmark gives an optimistic view of model capability.
- Overlap rates above 20% suggest the model may have memorised answers rather than generalised.
- Running this check before reporting results is good scientific hygiene.
# Check detailed_results.csv against UltraFeedback (5k sample)
python contamination_check.py \
--eval_file detailed_results.csv \
--threshold 0.2
# Use a different corpus or split
python contamination_check.py \
--eval_file detailed_results.csv \
--corpus openai/gsm8k \
--corpus_split train \
--corpus_sample 10000 \
--ngram_size 13=== Contamination Check ===
Eval file: detailed_results.csv
Corpus: HuggingFaceH4/ultrafeedback_binarized
N-gram size: 13
Threshold: 20%
Loaded 40 unique eval questions/IDs.
Index size: 2,847,193 unique 13-grams.
=== Results ===
Questions checked: 40
Contaminated (>20%): 2 (5.0%)
Avg overlap fraction: 0.0312
Flagged questions:
[3] overlap=0.241 preview='HumanEval/3'
[7] overlap=0.228 preview='HumanEval/7'
Saved report → results/contamination_report.json
Computes bootstrap confidence intervals (10,000 resamples) for all benchmark metrics and runs pairwise permutation tests between models.
# Full report: CIs + pairwise significance + power analysis
python statistical_significance.py --results detailed_results.csv
# Standalone power analysis
python statistical_significance.py --power --effect_size 0.02Metric Score 95% CI N
--------------------------------------------------------------------
allam-2-7b 10.0% [ 3.3%, 23.3%] 10
llama-3.1-8b-instant 50.0% [23.3%, 76.7%] 10
openai/gpt-oss-20b 10.0% [ 3.3%, 23.3%] 10
meta-llama/llama-4-scout-17b 100.0% [100.0%, 100.0%] 10
Or, for Qwen2.5-7B-Instruct on GSM8K + HumanEval:
| Metric | Score | 95% CI | N |
|---|---|---|---|
| GSM8K | 54.0% | [51.2%, 56.8%] | 500 |
| HumanEval | 70.0% | [67.3%, 72.6%] | 164 |
Comparison delta p-value Sig?
llama-3.1-8b-instant vs meta-llama/llama-4-scout -0.500 0.0120 *
allam-2-7b vs llama-3.1-8b-instant -0.400 0.0340 *
(* = p < 0.05, two-sided permutation test)
Use the --power flag to estimate how many eval examples are needed to reliably detect a given accuracy difference:
python statistical_significance.py --power --effect_size 0.02
# To detect a 2.0% accuracy difference with 80% power (alpha=0.05), you need ~2402 examples.Full table (80% power, alpha = 0.05):
| Effect size | Required N |
|---|---|
| 1% | 9,608 |
| 2% | 2,402 |
| 3% | 1,068 |
| 5% | 385 |
| 10% | 97 |
Takeaway: The default 10-problem HumanEval subset is only powered to detect differences of ~30 pp or larger. Use at least 385 problems for 5 pp sensitivity and 2,402 for 2 pp sensitivity.
Closes the eval→train→re-eval loop: failures from eval_tool_use.py become SFT training examples; the fine-tuned model is re-evaluated to measure the improvement.
python generate_sft_data.py \
--input results/tool_use_results.json \
--output data/sft_from_failures.jsonl \
--augment 3 # 3× rephrasings per failure
--real_data # also pull from Salesforce/xlam-function-calling-60k
--real_n 100Per-category correction strategy:
| Category | Strategy |
|---|---|
computation |
Reteach correct python_exec / calculator call |
error_recovery |
Add try/fallback pattern: fail → detect → correct call |
multi_step |
Explicit numbered sub-steps in CoT |
knowledge_compute |
lookup → store result → plug into calculator |
Output: JSONL with {"messages": [...], "source": "synthetic"/"real_api", "category": "..."}.
python finetune_on_failures.py \
--model Qwen/Qwen2.5-7B-Instruct \
--sft_data data/sft_from_failures.jsonl \
--output_dir checkpoints/tool_use_sft \
--output results/finetuning_results.json
# Dry-run (no GPU needed — prints config + step count)
python finetune_on_failures.py --dry_runUses QLoRA (4-bit NF4, LoRA r=16, α=32), 3 epochs, lr=2e-4. Prints a before/after comparison table per category.
Evaluates function-calling accuracy on real developer queries from Salesforce/xlam-function-calling-60k (60K examples across APIs).
| Metric | Description |
|---|---|
name_acc |
Correct function name called |
args_acc |
Correct arguments (normalised: lowercase strings, 2-decimal floats) |
full_acc |
Both name and args correct |
halluc_rate |
Fraction of calls naming a non-existent function |
# Demo mode (rule-based agent, no GPU)
python real_api_eval.py --mode demo --n 50
# Real model
CUDA_VISIBLE_DEVICES=0 python real_api_eval.py \
--model Qwen/Qwen2.5-7B-Instruct --n 200
# Side-by-side synthetic vs real
python real_api_eval.py --compare --n 100Evaluates structured output reliability — how accurately a model produces JSON that conforms to an explicit schema. Directly maps to production API usage where callers define schemas and expect strict adherence.
10 schemas: user_profile, api_response, tool_call, search_result, code_review, calendar_event, product_listing, error_response, model_completion, agent_action
Four error modes measured:
| Metric | Description |
|---|---|
strict_acc |
Valid JSON + all required fields + correct types (passes jsonschema.validate) |
partial_acc |
Valid JSON + required fields present but ≥1 type mismatch |
halluc_rate |
Extra keys not defined in the schema |
missing_rate |
Required fields absent |
Baseline — demo mode (correct outputs, validates harness):
| Schema | strict_acc | halluc_rate | missing_rate |
|---|---|---|---|
| user_profile | 100.0% | 0.0% | 0.0% |
| tool_call | 100.0% | 0.0% | 0.0% |
| agent_action | 100.0% | 0.0% | 0.0% |
| … (10 schemas) | 100.0% | 0.0% | 0.0% |
python eval_json_conformance.py --mode demo # verify harness (no GPU)
CUDA_VISIBLE_DEVICES=0 python eval_json_conformance.py \
--model Qwen/Qwen2.5-7B-Instruct --n 50 # real model evalExtends the single-turn ReAct harness to 3–5 turn conversations where each turn depends on the previous. Tests context retention, instruction fidelity over long conversations, and self-correction.
20 tasks across 4 categories:
| Category | Tasks | What it tests |
|---|---|---|
reference_chain |
5 | "Compute X" → "Multiply that by Y" → "What was my first answer?" |
state_tracking |
5 | Shopping cart / inventory / to-do list modified across turns |
instruction_stack |
5 | Each turn adds a new constraint; ALL must hold simultaneously by the final turn |
recovery_test |
5 | Injected wrong answer in turn 2; measures whether model self-corrects when asked to verify |
Three failure modes measured:
- Context loss — model ignores or contradicts a prior turn's result (
depends_oncheck) - Instruction drift — model drops accumulated constraints by turn 4–5
- Recovery — model self-corrects after an injected error
Measured results — demo ceiling (rule-based agent):
| Category | Success | Context loss | Drift | Recovery |
|---|---|---|---|---|
| reference_chain | 5/5 | 0 | 0 | — |
| state_tracking | 5/5 | 0 | 0 | — |
| instruction_stack | 5/5 | 0 | 0 | — |
| recovery_test | 5/5 | 0 | 0 | 5/5 |
| OVERALL | 20/20 (100%) | 0% | 0% | 100% |
python eval_multiturn.py --mode demo
CUDA_VISIBLE_DEVICES=0 python eval_multiturn.py --mode hf \
--model Qwen/Qwen2.5-7B-Instruct200 pairs of (correct instruction, adversarial paraphrase) across 5 strategies. Measures whether a model follows the original intent or is fooled by subtle rewording — the core thing post-training tries to fix.
5 adversarial strategies (40 pairs each):
| Strategy | Example | Risk |
|---|---|---|
format_override |
"Respond in JSON" + XML example appended | Example hijacks format |
late_contradiction |
Correct instruction early + contradiction appended at end | Late text overrides earlier |
implicit_negation |
"Don't NOT use bullet points" | Double negative confuses model |
example_hijack |
Correct rule + wrong-format demo example | Model imitates example over rule |
scope_creep |
"Answer in 1 sentence" + "cover all aspects" | Extra requirements make constraint impossible |
Measured results — demo agent (rule-based baseline):
| Strategy | Adherence | Fooling | Partial |
|---|---|---|---|
| format_override | 60.0% | 5.0% | 35.0% |
| late_contradiction | 90.0% | 10.0% | 0.0% |
| implicit_negation | 55.0% | 37.5% | 7.5% |
| example_hijack | 75.0% | 5.0% | 20.0% |
| scope_creep | 100.0% | 0.0% | 0.0% |
| OVERALL | 76.0% | 11.5% | 12.5% |
Key finding: implicit_negation is the most dangerous strategy — 37.5% fooling rate. implicit_negation + late_contradiction account for 83% of all failures.
python eval_contrastive_instructions.py --mode demo # no GPU
python eval_contrastive_instructions.py --mode demo --compare # side-by-side table
CUDA_VISIBLE_DEVICES=0 python eval_contrastive_instructions.py \
--model Qwen/Qwen2.5-7B-Instruct --n 100When building code agents or automation systems, accuracy alone is not enough. You also care about response time, token efficiency, cost, and reliability under tool-call failures. This project provides a simple, reproducible way to compare those tradeoffs — including a ReAct agent harness that tests error recovery, multi-step chaining, and tool accuracy, not just single-turn pass@1.
The failure-driven SFT loop (generate_sft_data.py → finetune_on_failures.py) closes the gap between evaluation and improvement: failures are not just diagnosed but corrected with targeted training data, making the benchmark actionable rather than just diagnostic.