Skip to content

Cache both harnesses under HF_HOME - #120

Merged
haideraltahan merged 4 commits into
mainfrom
harness-caches-under-hf-home
Oct 5, 2026
Merged

haideraltahan merged 4 commits into
mainfrom
harness-caches-under-hf-home

Conversation

@haideraltahan

@haideraltahan haideraltahan commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Neither harness caches under HF_HOME on its own.

lighteval hardcodes cache_dir = "~/.cache/huggingface/lighteval" and never reads HF_HOME, so every run writes into the user's home. It caches unconditionally — TransformersModel.__init__ constructs a SampleCache with no flag guarding it — so on LUMI (20 GB home) a run dies during model init:

transformers_model.py:238 in __init__ → SampleCache → cache_dir.mkdir()
OSError: [Errno 122] Disk quota exceeded:
'/users/<me>/.cache/huggingface/lighteval/EleutherAI/pythia-160m/e29cf52df2996268'

That cost one of four evaluations in a run I would otherwise have recorded as complete.

lm-eval keeps request caches in its own install directory, which is read-only in the container:

$ python -c "from lm_eval.caching.cache import PATH; print(PATH)"
/opt/conda/envs/py_3.12/lib/python3.12/site-packages/lm_eval/caching/.cache
$ touch .../lm_eval/caching/.probe
touch: Read-only file system

The change

export LIGHTEVAL_CACHE_DIR="$HF_HOME/lighteval"
export LM_HARNESS_CACHE_PATH="$HF_HOME/lm_eval"

lighteval only honours cache_dir as a model argument, so it is passed there in both the venv and container branches. LM_HARNESS_CACHE_PATH is read from the environment, so it is added to SINGULARITY_ENV_ARGS for the clusters that run with --contain (jupiter, snellius), which forward only the variables named there.

Verified on dev-g — lighteval now hits the shared path:

[CACHING] Loaded 4 cached indices for task 'flores200:deu_Latn-eng_Latn|4 …'
from /scratch/project_465002530/cache/lighteval/EleutherAI/pythia-160m/…/GENERATIVE.parquet

cache_dir exists on lighteval's ModelConfig in both 0.13.0 and 0.13.1.dev0.

lm-eval's --cache_requests stays opt-in

It is wired up behind CACHE_REQUESTS (true to enable, refresh to rebuild) rather than on by default, because enabling it makes limited runs several times slower. From lm_eval/api/task.py:

# process all documents when caching is specified for simplicity
if cache_requests and (not cached_instances or rewrite_requests_cache) and limit is not None:
    limit = None

--limit is ignored while building contexts, so a smoke test builds every document in the task. Same two tasks, same venv, same --limit 4, on dev-g:

task 0 task 1
without --cache_requests 2:56 3:25
with it, cold cache 11:17 11:11
with it, warm cache 10:38 13:53

The warm run is no faster than the cold one, and both are 3–4x the uncached baseline. The cache does work — the entry written by run 1 was not rewritten by run 2, which is the hit path — it just does not pay for itself here, and --limit is what the docs recommend for dev-g.

Two things to know if it is ever turned on by default:

  • The key is requests-{task}-{shots}shot-rank{r}-world_size{w}-tokenizer{name}. It does not cover the task YAML, so editing a task definition silently reuses the old prompts — hence CACHE_REQUESTS=refresh.
  • tokenizer_name is tokenizer.name_or_path, which for repo,revision=… is just the repo id, so every revision of a checkpoint repo shares one entry. Fine while the tokenizer is identical across revisions.

Tests

tests/test_harness_cache_paths.py — both paths exported under HF_HOME, cache_dir passed to lighteval in both branches, LM_HARNESS_CACHE_PATH forwarded into contained containers, and --cache_requests absent unless CACHE_REQUESTS is set. The first three fail on main. Suite passes (140), ruff check and format clean.

Merged main to pick up #110; the batching and caching changes touch the same invocation and both survive (test_task_batching.py passes).

lighteval hardcodes ~/.cache/huggingface/lighteval and never reads
HF_HOME, so every run fills the user's home directory. On LUMI that ends
as "OSError: [Errno 122] Disk quota exceeded" partway through an
evaluation, which the job then reports as finished.

lm-eval keeps request caches inside its own install directory, which is
a read-only filesystem in the container.

Export LIGHTEVAL_CACHE_DIR and LM_HARNESS_CACHE_PATH under HF_HOME, pass
lighteval its cache_dir as a model argument (the only way it honours
one), and forward LM_HARNESS_CACHE_PATH to clusters that run the
container with --contain.

lighteval caches unconditionally. lm-eval's --cache_requests stays
opt-in behind CACHE_REQUESTS, because it makes lm-eval ignore --limit
while building contexts.
@haideraltahan
haideraltahan force-pushed the harness-caches-under-hf-home branch from 0c13b47 to b6e4299 Compare September 22, 2026 15:21
@haideraltahan
haideraltahan merged commit e2204c3 into main Oct 5, 2026
3 checks passed
@haideraltahan
haideraltahan deleted the harness-caches-under-hf-home branch October 5, 2026 22:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants