chore: real token usage, Retry-After, mypy in CI, ruff 0.16 - #37
Conversation
- llm_client records provider-reported input/output tokens and models per document; ProcessingMetrics and BatchMetrics expose them. - Retries honour Retry-After (capped at 60 s) or back off exponentially with jitter. - mypy over src/docket runs in the CI lint job; 115 errors fixed. - ruff 0.16.8 (pinned in the dev extra); 201 findings fixed, C408 and DTZ ignored and FastAPI's File/Form defaults allowed, with reasons in config. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 71d2889292
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| seconds = None | ||
| if seconds is not None: | ||
| return min(max(seconds, 0.0), _MAX_RETRY_AFTER_S) | ||
| return _BACKOFF_S * (2 ** attempt) * random.uniform(0.5, 1.5) |
There was a problem hiding this comment.
Cap the exponential fallback delay
When DOCKET_LLM_RETRIES is set near its supported maximum of 10 and the provider omits Retry-After, this uncapped expression can sleep for as much as 1,536 seconds before the final retry and roughly 51 minutes cumulatively. That defeats the surrounding goal of avoiding workers being held for minutes and is a major regression from the previous linear schedule; apply a per-delay cap to the fallback path as well.
Useful? React with 👍 / 👎.
chore: real token usage, Retry-After, mypy in CI, ruff 0.16
The four smaller items from the external audit.
Summary
llm_clientreads what the provider reports — Ollamaprompt_eval_count/eval_count, OpenAI-compatibleusage.prompt_tokens/completion_tokens— per call. NewProcessingMetricsfields:llm_input_tokens,llm_output_tokens,llm_unreported_calls(calls with no counts, not in the sums),llm_models.BatchMetricssums the token fields. The chars/4 estimate is kept.Retry-After(seconds or HTTP date, capped at 60 s); without it, exponential backoff with jitter (2 s × 2^attempt × U(0.5, 1.5)).[tool.mypy]oversrc/docket; the lint job installs.[dev,review]and runsmypy. 115 errors fixed — annotations, narrowing, and loop-variable reuse; none turned out to be a runtime bug (the_turn(*c, ...)and duplicateby_rowfindings were false positives, now written so mypy can see it).devextra. 201 findings: 160 auto-fixed (import order, unused noqa, quoted annotations, …), the rest by hand. Config ignoresC408(kwargs-styledict()in factories) andDTZ(document dates are calendar dates), and allows FastAPI'sFile/Form/Depends/Headerdefaults. Intentional broad excepts carry anoqawith the reason.Union[...]→X | Yspelling.Test plan
🤖 Generated with Claude Code