Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
50a7d15
Harden unified evaluator against candidate tampering
ahydchh Sep 7, 2026
1623c6e
Add candidate sandbox helper and stop trusting self-reported scores
ahydchh Sep 7, 2026
f7b1375
Isolate joint_replenishment candidate and validate its inputs
ahydchh Sep 7, 2026
660752e
InventoryOptimization: run candidates in a subprocess, validate their…
ahydchh Sep 7, 2026
4439c61
JobShop: take the instance from the task, not from the candidate
ahydchh Sep 7, 2026
fee14fa
EngDesign: isolate the submission and recompute its reported metrics
ahydchh Sep 7, 2026
769f6f1
Optics: move the forward model and the ruler to the scorer (adaptive,…
ahydchh Sep 7, 2026
cf52dd2
MolecularMechanics: give each eval stage its own wall-clock cap
ahydchh Sep 7, 2026
eb6fe5e
InventoryOptimization: bound the evaluator stage in run_eval.py
ahydchh Sep 7, 2026
006d816
QuantumComputing: check the optimized circuit is equivalent before sc…
ahydchh Sep 7, 2026
0bc0306
EnergyStorage: run the charging policy out of process, reject non-fin…
ahydchh Sep 7, 2026
9fb6533
PyPortfolioOpt: make the risk constraints a gate, not a penalty term
ahydchh Sep 7, 2026
1db7c36
MallocLab: take the score off the channel the candidate can write to
ahydchh Sep 7, 2026
cd97771
Optics holographic: the candidate submits a phase map, not an optical…
ahydchh Sep 7, 2026
2a3b69b
DiffSim, EV2Gym, CoFlyers, SustainDC: candidate out of process, metri…
ahydchh Sep 7, 2026
b187a6c
candidate_sandbox: document what process isolation does not buy
ahydchh Sep 7, 2026
a481978
MallocLab test: reimplement the old parser instead of fetching it fro…
ahydchh Sep 7, 2026
0c3b048
Detect a candidate poisoning the source benchmark tree, not just the …
ahydchh Sep 7, 2026
922bf06
UAV inspection, obstacle avoidance: own the environment and the scorer
ahydchh Sep 7, 2026
260d761
Fix the pyportfolioopt sandbox test to set the env the harness sets
ahydchh Sep 7, 2026
8f2dce2
StructuralOptimization: import before the candidate runs, and honour …
ahydchh Sep 7, 2026
28235fd
Lunar, predict_modality, Muon, HRS, CarAero: load the scorer before t…
ahydchh Sep 7, 2026
7046b99
sampler_isolation: the candidate could forge the clock its score depe…
ahydchh Sep 7, 2026
d182012
LDPC, PMD, Rayleigh: run the sampler in a subprocess, aggregate in th…
ahydchh Sep 7, 2026
c339607
candidate_sandbox: kill the candidate's whole process group, and stop…
ahydchh Sep 7, 2026
45d3c17
KernelEngineering: three processes, and every timed rep gets checked
ahydchh Sep 7, 2026
181f1ac
PIDTuning, RobotArm, Quadruped: pin the world model before the candid…
ahydchh Sep 7, 2026
afbc0f5
Cryptographic: compile once, before the candidate runs, and check eve…
ahydchh Sep 7, 2026
15520c2
Stop the robotics_b attack tests littering the shared temp root
ahydchh Sep 7, 2026
056b00d
Report tracked benchmark files the test suite leaves modified
ahydchh Sep 7, 2026
0e7e861
Stop handing the scorer's own solution to the solver as prompt context
ahydchh Sep 7, 2026
5502530
UAV: regenerate result_log, which recorded a score from a retired for…
ahydchh Sep 7, 2026
6feaa19
Fix evaluator isolation regressions and refresh corrected leaderboard
ahydchh Sep 14, 2026
7a958aa
Restore measured baselines and limit leaderboard documentation changes
ahydchh Sep 14, 2026
3849b27
Trim audit artifacts and local-only tests
ahydchh Sep 14, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
16 changes: 8 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,14 +127,14 @@ Detailed leaderboard (incl. average rank): [lab.einsia.ai/frontier-eng/leaderboa

| Rank | Model | Medal (v1) | Medal (v1-lite) | 🥇 | 🥈 | 🥉 |
| :--: | :--- | --: | --: | --: | --: | --: |
| 1 | GPT-5.4 | 0.596 | 0.667 | 24 | 5 | 2 |
| 2 | Claude Opus 4.6 | 0.490 | 0.501 | 9 | 18 | 6 |
| 3 | GLM-5 | 0.312 | 0.233 | 4 | 10 | 12 |
| 4 | DeepSeek V3.2 | 0.248 | 0.166 | 3 | 9 | 8 |
| 5 | Gemini 3.1 Pro Preview | 0.213 | 0.200 | 3 | 6 | 9 |
| 6 | Seed 2.0 Pro | 0.185 | 0.100 | 3 | 7 | 3 |
| 7 | Grok 4.20 | 0.184 | 0.133 | 3 | 6 | 5 |
| 8 | Qwen3 Coder Next | 0.121 | 0.000 | 3 | 3 | 2 |
| 1 | Claude Opus 4.6 | 0.533 | 0.501 | 14 | 15 | 3 |
| 2 | GPT-5.4 | 0.454 | 0.267 | 18 | 4 | 2 |
| 3 | GLM-5 | 0.347 | 0.300 | 7 | 8 | 12 |
| 4 | Gemini 3.1 Pro Preview | 0.277 | 0.267 | 7 | 7 | 4 |
| 5 | DeepSeek V3.2 | 0.269 | 0.299 | 6 | 6 | 8 |
| 6 | Grok 4.20 | 0.227 | 0.200 | 6 | 5 | 4 |
| 7 | Seed 2.0 Pro | 0.206 | 0.100 | 6 | 4 | 3 |
| 8 | Qwen3 Coder Next | 0.170 | 0.066 | 5 | 3 | 3 |

## Contributing

Expand Down
16 changes: 8 additions & 8 deletions README_zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -122,14 +122,14 @@ bash scripts/batch/validate_v1_task_envs.sh

| 排名 | Model | Medal (v1) | Medal (v1-lite) | 🥇 | 🥈 | 🥉 |
| :--: | :--- | --: | --: | --: | --: | --: |
| 1 | GPT-5.4 | 0.596 | 0.667 | 24 | 5 | 2 |
| 2 | Claude Opus 4.6 | 0.490 | 0.501 | 9 | 18 | 6 |
| 3 | GLM-5 | 0.312 | 0.233 | 4 | 10 | 12 |
| 4 | DeepSeek V3.2 | 0.248 | 0.166 | 3 | 9 | 8 |
| 5 | Gemini 3.1 Pro Preview | 0.213 | 0.200 | 3 | 6 | 9 |
| 6 | Seed 2.0 Pro | 0.185 | 0.100 | 3 | 7 | 3 |
| 7 | Grok 4.20 | 0.184 | 0.133 | 3 | 6 | 5 |
| 8 | Qwen3 Coder Next | 0.121 | 0.000 | 3 | 3 | 2 |
| 1 | Claude Opus 4.6 | 0.533 | 0.501 | 14 | 15 | 3 |
| 2 | GPT-5.4 | 0.454 | 0.267 | 18 | 4 | 2 |
| 3 | GLM-5 | 0.347 | 0.300 | 7 | 8 | 12 |
| 4 | Gemini 3.1 Pro Preview | 0.277 | 0.267 | 7 | 7 | 4 |
| 5 | DeepSeek V3.2 | 0.269 | 0.299 | 6 | 6 | 8 |
| 6 | Grok 4.20 | 0.227 | 0.200 | 6 | 5 | 4 |
| 7 | Seed 2.0 Pro | 0.206 | 0.100 | 6 | 4 | 3 |
| 8 | Qwen3 Coder Next | 0.170 | 0.066 | 5 | 3 | 3 |

## 贡献

Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,5 @@
# EVOLVE-BLOCK-START
from __future__ import annotations
import time

from qiskit import transpile
from qiskit.circuit import QuantumCircuit
Expand All @@ -13,143 +12,31 @@ def _cost(qc: QuantumCircuit) -> float:
return sum(inst.operation.num_qubits == 2 for inst in qc.data) + 0.2 * qc.depth()


def _post_optimize(qc: QuantumCircuit, target: Target) -> QuantumCircuit:
try:
from qiskit.transpiler import PassManager
from qiskit.transpiler.passes import (
Optimize1qGatesDecomposition, CXCancellation,
CommutativeCancellation, CommutationAnalysis,
)
pm = PassManager([
CommutationAnalysis(), CommutativeCancellation(),
CXCancellation(), Optimize1qGatesDecomposition(target=target),
])
return pm.run(qc)
except Exception:
return qc


def optimize_circuit(input_circuit: QuantumCircuit, target: Target, case: dict) -> QuantumCircuit:
qc_rewritten = optimize_by_local_rewrite(input_circuit)
"""Target-aware transpile search baseline for routing-heavy circuits."""
qc = optimize_by_local_rewrite(input_circuit)
if target is None:
return qc_rewritten

time_limit = 92
t0 = time.monotonic()
elapsed = lambda: time.monotonic() - t0

best = None
best_score = float('inf')
top_k = []
TOP_K = 8

def _update(cand):
nonlocal best, best_score, top_k
s = _cost(cand)
if s < best_score:
best = cand
best_score = s
if len(top_k) < TOP_K:
top_k.append((s, cand))
top_k.sort(key=lambda x: x[0])
elif s < top_k[-1][0]:
top_k[-1] = (s, cand)
top_k.sort(key=lambda x: x[0])

source_circuits = [input_circuit, qc_rewritten]

for pre_opt in (0, 1, 2):
for pre_seed in (0, 7, 42, 99, 137, 200):
if elapsed() > time_limit * 0.10:
break
try:
qc_pre = transpile(input_circuit, target=target,
optimization_level=pre_opt, seed_transpiler=pre_seed)
source_circuits.append(qc_pre)
_update(qc_pre)
except Exception:
pass

for approx in (0.99, 0.999):
for pre_seed in (0, 42):
if elapsed() > time_limit * 0.12:
break
try:
_update(transpile(input_circuit, target=target, optimization_level=3,
seed_transpiler=pre_seed, approximation_degree=approx))
except Exception:
pass
return qc

num_qubits = case.get("num_qubits", input_circuit.num_qubits)
best = qc
best_score = _cost(qc)
option_sets = (
{"optimization_level": 3, "layout_method": "sabre", "routing_method": "sabre"},
{"optimization_level": 3, "layout_method": "lookahead", "routing_method": "sabre"},
{"optimization_level": 3, "layout_method": "dense", "routing_method": "sabre"},
{"optimization_level": 3, "layout_method": "trivial", "routing_method": "sabre"},
{"optimization_level": 3},
{"optimization_level": 2, "layout_method": "sabre", "routing_method": "sabre"},
{"optimization_level": 2, "layout_method": "dense", "routing_method": "sabre"},
{"optimization_level": 2},
)

p1_end = time_limit * 0.48
for qc in source_circuits:
if elapsed() > p1_end:
break
for seed in range(600):
if elapsed() > p1_end:
break
for kw in option_sets:
try:
_update(transpile(qc, target=target, seed_transpiler=seed, **kw))
except Exception:
pass

if elapsed() < time_limit * 0.52:
for s, cand in list(top_k):
for seed in (num_qubits + 5, num_qubits + 11, num_qubits + 17, num_qubits + 29):
for transpile_kwargs in option_sets:
try:
_update(_post_optimize(cand, target))
candidate = transpile(qc, target=target, seed_transpiler=seed, **transpile_kwargs)
except Exception:
pass

p2_end = time_limit * 0.92
reopt_opts = (
{"optimization_level": 3, "layout_method": "sabre", "routing_method": "sabre"},
{"optimization_level": 3},
{"optimization_level": 2, "layout_method": "sabre", "routing_method": "sabre"},
{"optimization_level": 3, "layout_method": "dense", "routing_method": "sabre"},
)

for _round in range(25):
if elapsed() > p2_end:
break
improved = False
for _, cand in list(top_k):
if elapsed() > p2_end:
break
for seed in range(400):
if elapsed() > p2_end:
break
for ro in reopt_opts:
try:
old = best_score
_update(transpile(cand, target=target, seed_transpiler=seed, **ro))
if best_score < old:
improved = True
try:
_update(_post_optimize(best, target))
except Exception:
pass
except Exception:
pass
if not improved:
break

if best is not None and elapsed() < time_limit * 0.98:
for seed in range(500):
if elapsed() > time_limit * 0.98:
break
try:
_update(transpile(best, target=target, seed_transpiler=seed, optimization_level=3))
except Exception:
pass

return best if best is not None else qc_rewritten
continue
score = _cost(candidate)
if score < best_score:
best = candidate
best_score = score
return best
# EVOLVE-BLOCK-END
Original file line number Diff line number Diff line change
Expand Up @@ -20,12 +20,50 @@ def optimize_circuit(input_circuit: QuantumCircuit, target: Target, case: dict)
elif "rigetti" in target_name:
rigetti.add_equivalences(SessionEquivalenceLibrary)

kw = {"circuits": optimized, "target": target, "optimization_level": 3, "seed_transpiler": 42}
transpile_kwargs = {
"circuits": optimized,
"target": target,
"optimization_level": case.get("optimization_level", 3),
"seed_transpiler": 42,
}
if "ionq" in target_name:
kw["basis_gates"] = ["rz", "sx", "x", "rzz", "measure"]
transpile_kwargs["basis_gates"] = ["rz", "sx", "x", "rzz", "measure"]
if "ibm" in target_name or "rigetti" in target_name:
kw.update({"layout_method": "sabre", "routing_method": "sabre", "approximation_degree": 0.85, "unitary_synthesis_method": "sk", "unitary_synthesis_plugin_config": {"optimization_level": 3}})
# No `approximation_degree` here on purpose. Lowering it buys a smaller
# two-qubit count (247 -> 214 on case 01) by throwing away fidelity
# (0.23 against the input circuit), and the evaluator's equivalence
# gate rejects the result outright.
transpile_kwargs.update(
{
"layout_method": "sabre",
"routing_method": "sabre",
}
)

transpiled = transpile(**kw)
return optimize_by_local_rewrite(transpiled, max_rounds=32)
best_circuit = None
best_score = float('inf')

# Try multiple transpiler seeds to find a better layout/routing
for seed in [42, 123, 456, 789]:
transpile_kwargs["seed_transpiler"] = seed
try:
transpiled = transpile(**transpile_kwargs)
optimized_out = optimize_by_local_rewrite(transpiled, max_rounds=32)

# Score primarily by 2-qubit gate count, with depth as a tie-breaker
score = optimized_out.num_nonlocal_gates() * 10000 + optimized_out.depth()

if score < best_score:
best_score = score
best_circuit = optimized_out
except Exception:
# Fallback gracefully if a particular seed raises an issue
pass

if best_circuit is None:
transpile_kwargs["seed_transpiler"] = 42
transpiled = transpile(**transpile_kwargs)
best_circuit = optimize_by_local_rewrite(transpiled, max_rounds=32)

return best_circuit
# EVOLVE-BLOCK-END
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,30 @@ def _cost(qc: QuantumCircuit) -> float:
return sum(inst.operation.num_qubits == 2 for inst in qc.data) + 0.2 * qc.depth()


def _score(qc: QuantumCircuit, target: Target) -> tuple[float, QuantumCircuit]:
qc = transpile(optimize_by_local_rewrite(qc), target=target, optimization_level=0, seed_transpiler=10)
return _cost(qc), qc


def optimize_circuit(input_circuit: QuantumCircuit, target: Target, case: dict) -> QuantumCircuit:
"""Minimize the evaluator's routed cost with the cheapest valid circuit."""
qc = optimize_by_local_rewrite(input_circuit)
if target is None:
return optimize_by_local_rewrite(input_circuit)
return QuantumCircuit(*input_circuit.qregs, *input_circuit.cregs)
return qc
n = case.get("num_qubits", input_circuit.num_qubits)
try:
best = transpile(qc, target=target, optimization_level=1)
best_score = _cost(best)
except Exception:
best, best_score = qc, 1e18
opts = (
{"optimization_level": 3, "layout_method": "sabre", "routing_method": "sabre"},
{"optimization_level": 3, "layout_method": "dense", "routing_method": "sabre"},
{"optimization_level": 3, "layout_method": "lookahead", "routing_method": "sabre"},
{"optimization_level": 3},
)
for s in (n + 1, n + 7, n + 19, n + 31, 13):
for kw in opts:
try:
cand = transpile(qc, target=target, seed_transpiler=s, **kw)
except Exception:
continue
score = _cost(cand)
if score < best_score:
best, best_score = cand, score
return best
# EVOLVE-BLOCK-END
Loading
Loading