Repository navigation
WIP: Judge tuning (2/6): Add NePS configuration and sessions - #135
Draft
ErlisLushtaku wants to merge 12 commits into
Draft
ErlisLushtaku wants to merge 12 commits into
ErlisLushtaku wants to merge 12 commits into
Conversation
This was referenced Oct 7, 2026
…lislushtaku/feat/judge-tuning-02-neps-config
…lislushtaku/feat/judge-tuning-02-neps-config
…lislushtaku/feat/judge-tuning-02-neps-config
ErlisLushtaku
commented
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WIP. Not ready for review yet.
Adds a packaged, dataset-free tune-judge task with overridable target, objectives, fidelity, and optimizer defaults. Nested search spaces use lists or ranges and value priors from the resolved reference judge. Numeric domains follow the target config field, so integer bounds on temperature still produce a continuous parameter.
The target meta-eval task supplies judge defaults and restrictions. Explicit run directories support resume and helper workers while configuration hashes protect scientific settings. Dataset downloads skip tuning tasks.
Part 2 of 6, stacked on #134. Native usage follows in #140, execution in #133, reference pricing in #139, then docs in #136. All remain WIP.