Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion TASK_DETAILS.md
Original file line number Diff line number Diff line change
Expand Up @@ -204,14 +204,18 @@ We welcome new engineering problem ideas — even without complete verification
<td>Polarization-multiplexed holography</td>
</tr>
<tr>
<td rowspan="2"><b>ComputerSystems</b></td>
<td rowspan="3"><b>ComputerSystems</b></td>
<td><code>MallocLab</code></td>
<td>High-performance C memory allocator (utilization &amp; throughput)</td>
</tr>
<tr>
<td><code>DuckDBWorkloadOptimization</code></td>
<td>Index / materialized-view selection and query rewriting on official DuckDB workloads</td>
</tr>
<tr>
<td><code>CacheReplacementPolicyOptimization</code></td>
<td>Cache eviction policy optimization on real KV-cache request traces</td>
</tr>
<tr>
<td><b>EngDesign</b></td>
<td><code>CY_03, WJ_01, XY_05, AM_02, AM_03, YJ_02, YJ_03</code></td>
Expand Down
6 changes: 5 additions & 1 deletion TASK_DETAILS_zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -204,14 +204,18 @@ Frontier-Eng 目前已覆盖以下领域的任务。每个任务均配有可运
<td>偏振复用全息</td>
</tr>
<tr>
<td rowspan="2"><b>ComputerSystems</b></td>
<td rowspan="3"><b>ComputerSystems</b></td>
<td><code>MallocLab</code></td>
<td>高性能 C 动态内存分配器(utilization &amp; throughput)</td>
</tr>
<tr>
<td><code>DuckDBWorkloadOptimization</code></td>
<td>基于 DuckDB 官方 workload 的索引 / 物化视图选择与查询改写</td>
</tr>
<tr>
<td><code>CacheReplacementPolicyOptimization</code></td>
<td>基于真实 KV-cache 请求 trace 的缓存淘汰策略优化</td>
</tr>
<tr>
<td><b>EngDesign</b></td>
<td><code>CY_03, WJ_01, XY_05, AM_02, AM_03, YJ_02, YJ_03</code></td>
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
# The two trace files are sha256-checked by verification/evaluator.py.
# Store and check them out verbatim (no end-of-line conversion) so the
# digests recorded in references/constants.json hold on every platform.
*.csv -text
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Keep the frozen real trace data. The repository root .gitignore excludes
# `*.csv` (whitelisting only `leaderboard/*.csv`), but both traces below are
# sha256-checked by verification/evaluator.py and must ship verbatim.
!traces/*.csv
!verification/heldout/*.csv

__pycache__/
*.pyc
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# Cache Replacement Policy Optimization

A Frontier-Eng benchmark: an **online cache replacement (eviction) policy** is
optimised against a **real Meta production kvcache trace** published by CacheLib
for CacheBench, and scored by **hit rate** at a fixed cache capacity, compared
against CacheLib's default LRU.

## Layout

| path | role |
| --- | --- |
| `policy.py` | the agent-editable submission (initial program: LRU, `FE-BATCH-JSON-STDIO-V1`) |
| `traces/kvcache_202206_traces_1.csv` | the agent-visible real trace (a verbatim prefix of CacheLib's published `pub/kvcache/202206/kvcache_traces_1.csv`) |
| `verification/evaluator.py` | frozen evaluator: replays the submission on the held-out trace and prints `{"valid":..., "combined_score": hit_rate, "metrics": {...}}` |
| `verification/reference_cache.py` | frozen replay model + reference policies (LRU / FIFO / LRU-2Q) + the offline optimum (Belady 1966) |
| `verification/heldout/` | the evaluator-only real trace (a different published file; never in the agent's manifests) |
| `baseline/solution.py` | frozen baseline = CacheLib's default LRU |
| `references/constants.json` | frozen numbers, sources and the score definition |
| `docs/PROVENANCE.md` | where every number comes from (CacheLib sources, workload, capacity, metric) |
| `docs/agent_notes.md` | notes for the agent (model, protocol, local measurement) |
| `frontier_eval/` | the canonical execution manifests |

## Run it locally

```bash
# the frozen baseline over the agent-visible trace
python - <<'PY'
import json, subprocess, sys, tempfile, pathlib
sys.path.insert(0, "verification")
import reference_cache as R
keys, stats = R.load_reference_stream("traces/kvcache_202206_traces_1.csv")
cap = 422
instance = {"protocol": "FE-BATCH-JSON-STDIO-V1",
"runs": [{"name": "dev", "cache_size": cap, "keys": keys}]}
with tempfile.TemporaryDirectory() as d:
p = pathlib.Path(d) / "i.json"; p.write_text(json.dumps(instance))
out = subprocess.run([sys.executable, "baseline/solution.py", str(p)], capture_output=True, text=True)
print("baseline hit rate:", R.replay_evictions(keys, cap, json.loads(out.stdout)["runs"][0]["evictions"])["hit_rate"])
PY
```

## Provenance

Every constant is derived from a real CacheLib artifact — the CacheBench
documentation, the CacheBench configuration schema and the Meta trace files
CacheLib publishes — and the derivation is written down per item in
`docs/PROVENANCE.md`, including what was **rejected** as unsupported (composite
score weights, invented cache sizes, LLM-generated workloads).

## Frontier-Eval onboarding

* unified benchmark id: `ComputerSystems/CacheReplacementPolicyOptimization`
* editable program: `policy.py` (only the `EVOLVE-BLOCK` region is scored)
* validation: `python verification/evaluator.py policy.py`
* structured result: the unified `eval_command` writes the evaluator's single stdout
JSON object to `metrics.json` in the evaluation cwd, which is the file `task=unified`
reads (`parse_stdout_json` is disabled repo-wide, so the evaluator's own stdout is
not consumed by the framework)
* runtime overrides: **none** — Python standard library only; no network, no
Docker and no third-party or system toolchain requirement
(see `verification/requirements.txt`)
127 changes: 127 additions & 0 deletions benchmarks/ComputerSystems/CacheReplacementPolicyOptimization/Task.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
# Cache Replacement Policy Optimization (real Meta kvcache trace)

## Background

Production key-value caches (memcached/Meta kvcache-style, managed by CacheLib)
are capacity-bound: the DRAM budget is fixed, the key population is far larger
than the cache, and every object that is not resident costs a backend lookup.
CacheLib's own benchmark tool, **CacheBench**, is built around exactly this
question — its documentation describes its first use case as

> "Prototype and evaluation of cache heuristics: CacheBench can be used to
> compare the cache performance (**hit ratio**) of various configuration options
> for existing heuristics and new heuristics. For example, given a cache size
> and workload, comparing LRU vs 2Q vs FIFO vs new heuristic added to
> CacheLib."

This task is that experiment: the same real trace, the same cache capacity, the
same request order — only the **eviction policy** changes.

## Engineering Value

At Meta's cache scale a few points of hit ratio translate into a
proportional cut in backend QPS, network traffic and tail latency, without any
extra DRAM. Replacement policy is the classic place where that gain is won:
SIEVE (NSDI'24), S3-FIFO (SOSP'23), ARC (FAST'03) and 2Q are all published
eviction policies whose entire claim is a lower miss ratio than LRU on real
production traces at a fixed cache size. A service that moves from LRU to a
better policy gets the improvement on every host it runs on.

## Objective

Given one candidate program (`policy.py`), implement an **online cache
replacement policy**: for a full-associativity cache of fixed capacity C, decide
which resident object to evict when a lookup misses and the cache is full.
Because every lookup on a missing object loads that object, the eviction
decision at each miss is what determines the hit rate.

Your policy is scored on **hit rate on a real production trace you never
receive** at a fixed capacity. You *do* get one real trace (see `README.md`) to
develop and measure against locally.

## Submission Contract

`policy.py` is the submission. Only the region between the EVOLVE-BLOCK markers
may be edited:

```python
class Policy:
def __init__(self, capacity): ...
def access(self, key):
"""Return the key to evict on this access, or None if nothing is evicted."""
```

**Runtime execution contract** (`FE-BATCH-JSON-STDIO-V1`): the frozen evaluator
runs

```
python policy.py <instance.json> # stdin is CLOSED
```

and your program must print **exactly one JSON object on stdout** and exit 0:

```json
{"runs": [{"name": "<the run name>", "evictions": ["<key>", null, ...]}]}
```

* one entry per element of `instance["runs"]`, in that order, with the same
`name`;
* `evictions[i]` is the key you evict at access `i` of that run, or `null` when
nothing is evicted;
* a hit must never evict; a miss on a full cache must name one **resident** key;
* the evaluator does the replay itself and counts the hits — your program never
reports hits or misses, so it cannot report a hit rate.

The instance carries `runs[i]["cache_size"]` (the fixed capacity) and
`runs[i]["keys"]` (the real request stream, in replay order).

## What is fixed (and checked)

* the workload, the cache capacity, the reference model, the evaluator, the
baseline and `frontier_eval/constraints.txt` — all frozen;
* the graded trace is a *different* real trace of the same family, delivered to
your program only as the `keys` list of the instance at run time;
* your policy must be **online**: the evaluator replays you again on the first
half of the same run with the same capacity; an online policy makes exactly
the same decisions, so reading ahead (an offline/Belady-style oracle) is
detected and rejected;
* your hit rate may not exceed the offline optimum (Belady 1966), which is the
literature's reference bound and is not reachable by a legal online policy.

## Scoring

`combined_score` is the **hit rate** on the held-out trace at the fixed
capacity: `hits / lookups`. No weights, no bonuses, no penalties. For reference
the same evaluator reports the frozen LRU baseline hit rate and the offline
optimum, and the evaluator's own reference replays are: LRU 0.2663, FIFO 0.2640,
LRU-2Q 0.2835, offline optimum 0.4325 at the graded capacity. The baseline
(LRU) is a legal, valid submission — beating it is where the score comes from.

Rejected submissions (protocol, legality, causality, or exceeding the optimum)
report `combined_score = -1e18`, the frozen "no score" sentinel.

## Evaluator output (the frozen reporting contract)

The evaluator prints exactly one JSON object on stdout and always exits 0:

```json
{"valid": <bool>, "combined_score": <float>, "metrics": {...}}
```

`valid` means "legal online policy, graded and reproducible" — it is not
"good": the quality is `combined_score`. A rejected submission also carries
`error`, and `combined_score` is `-1e18`. A fully graded submission reports
exactly these `metrics` keys — integer counts:

`capacity_objects`, `lookups`, `working_set_objects`, `hits`, `misses`,
`causality_probe_lookups`

and floats:

`capacity_fraction`, `baseline_hit_rate_lru`, `baseline_hit_rate_fifo`,
`baseline_hit_rate_lru_2q`, `offline_optimum_hit_rate`, `hit_rate`,
`miss_ratio`, `improvement_over_lru`, `fraction_of_offline_optimum`,
`candidate_runtime_seconds`, `causality_probe_runtime_seconds`

Only the last two are measured wall-clock timings; every other reported value
is a frozen constant or a deterministic function of the submission's decisions.
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
{
"baseline": "LRU (CacheLib LruAllocator default)",
"command": "{python} baseline/solution.py {problem}",
"graded_workload": {
"combined_score": 0.2662958338063344,
"command": "{python} verification/evaluator.py {candidate} {problem}",
"metrics": {
"baseline_hit_rate_lru": 0.2662958338063344,
"capacity_objects": 417,
"hit_rate": 0.2662958338063344,
"hits": 11729,
"lookups": 44045,
"misses": 32316,
"offline_optimum_hit_rate": 0.43246679532296517
},
"valid": true
},
"reference_policies_on_the_graded_workload_at_the_graded_capacity": {
"fifo_hit_rate": 0.26398002043364743,
"lru_2q_hit_rate": 0.28348280168009987,
"lru_hit_rate": 0.2662958338063344,
"offline_optimum_hit_rate": 0.43246679532296517
},
"visible_workload": {
"cache_size": 422,
"distinct_keys": 21143,
"exit_code": 0,
"hit_rate": 0.264167254362,
"hits": 11612,
"illegal": null,
"lookups": 43957,
"miss_ratio": 0.735832745638,
"misses": 32345,
"runtime_seconds": 0.129,
"sha256": "5503257dbd9d73c455754c6caebef8409cecf17188dfef2cad195acfdb8715d1",
"trace": "traces/kvcache_202206_traces_1.csv"
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
CacheReplacementPolicyOptimization baseline
policy : LRU (CacheLib LruAllocator default)
command : {python} baseline/solution.py {problem}

visible workload : traces/kvcache_202206_traces_1.csv (real Meta kvcache trace, CacheLib published)
sha256 5503257dbd9d73c455754c6caebef8409cecf17188dfef2cad195acfdb8715d1
lookups=43957 distinct_keys=21143 cache_size=422
exits 0, hits=11612 misses=32345 hit_rate=0.264167254362 miss_ratio=0.735832745638
illegal_evictions=None, runtime=0.129s

graded workload : evaluator-only real trace (see docs/PROVENANCE.md section 2)
evaluator: {python} verification/evaluator.py {candidate} {problem}
valid=True combined_score=0.266295833806 (hit rate)

reference replays on the graded workload at the graded capacity (frozen):
LRU 0.266296 (hits 11729)
FIFO 0.263980 (hits 11627)
LRU-2Q 0.283483 (hits 12486)
offline opt 0.432467 (hits 19048) <- bound only, never a submission
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
#!/usr/bin/env python3
# CANDIDATE-PROTOCOL: FE-BATCH-JSON-STDIO-V1
"""baseline/solution.py -- FROZEN baseline (do not edit).

CLI: python baseline/solution.py <instance.json> (stdin is CLOSED)
prints exactly one JSON object: {"runs": [{"name":..., "evictions":[...]}]}

The baseline is CacheLib's default DRAM eviction policy, LRU
(`LruAllocator`). CacheBench's documented way to compare replacement
heuristics is exactly this: "given a cache size and workload, comparing LRU vs
2Q vs FIFO vs new heuristic added to CacheLib" (CacheLib CacheBench overview).
The evaluator replays this same baseline over the held-out workload; the
frozen hit counts are recorded in references/constants.json.

Only the standard library is used.
"""

import json
import sys
from collections import OrderedDict


class Lru:
def __init__(self, capacity):
self.capacity = int(capacity)
self.resident = OrderedDict()

def access(self, key):
if key in self.resident:
self.resident.move_to_end(key)
return None
if len(self.resident) < self.capacity:
self.resident[key] = True
return None
victim, _ = self.resident.popitem(last=False)
self.resident[key] = True
return victim


def solve(instance):
runs = []
for run in instance["runs"]:
policy = Lru(run["cache_size"])
runs.append({
"name": run["name"],
"evictions": [policy.access(key) for key in run["keys"]],
})
return {"runs": runs}


def main():
with open(sys.argv[1], "r", encoding="utf-8") as handle:
instance = json.load(handle)
print(json.dumps(solve(instance)))


if __name__ == "__main__":
main()
Loading
Loading