Skip to content

🏷️ fix: Scope Tool-Output Redaction to Tool Calls, Not Label Evidence - #574

Merged
danny-avila merged 1 commit into
mainfrom
fix/label-trace-only-redaction
Sep 28, 2026
Merged

danny-avila merged 1 commit into
mainfrom
fix/label-trace-only-redaction

Conversation

@danny-avila

@danny-avila danny-avila commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

What breaks

A Langfuse tool-output redaction policy (toolOutputTracing.redactedToolNames or enabled: false) exists to redact the named tools' outputs on their tool-call observations. The three label calls also apply it to the evidence they send the label model, and with any active policy they drop every piece of free-form evidence besides:

Call Under an active policy
generateActivityLabel Matching tool outputs replaced; reasoning, intent and previous headers dropped; a reasoning-only block returns {} without a model call.
generateActivityPhaseLabel Child labels, reasoning and assistant commentary dropped. A phase of labeled activities, the normal case once hosts replace entries with committed labels, builds an empty prompt and returns {}.
generateReasoningLabel buildReasoningLabelPrompt returns '', so no reasoning label is produced.

A deployment that redacts one tool it rarely calls therefore loses every phase title and reasoning label and gets thinner activity labels, on every run and whether or not Langfuse is exporting.

Change

The label calls no longer resolve or apply the tool-output redaction policy. Each builds its prompt from the full evidence:

/** Tool-output redaction applies to tool-call observations, not to the
 *  label model's evidence, so the label is written from the full batch. */
const userPrompt = buildActivityLabelPrompt({
  entries,
  charLimit,
  thinkingExcerpts,
  lastAssistantText:
    lastAssistantPhase === 'final_answer' ? undefined : lastAssistantText,
  previousLabels,
});

Removed with it:

  • the per-call policy resolution and the multi-agent union in all three calls
  • the freeFormSuppressed early return in generateActivityLabel
  • the unattributed / omitted-activity bookkeeping in generateActivityPhaseLabel, which only chose the policies to union

A call still returns {} when the full evidence produces an empty prompt.

What stays redacted

Tool-call observations. src/langfuse.ts and the toolOutputTracing pipeline are unchanged: matching tools' outputs are still replaced on their tool spans, and tool-message values in traced generation inputs are still redacted by the existing value walker. The Langfuse-callback tests that assert this through the real export path (for example "promotes trusted ToolMessage artifact metadata before redacting callback output") pass unchanged.

What changes for hosts

Label calls are traced as their own generations. Their traced input is the prompt the label model received, so it now contains the tool outputs, reasoning and commentary the label was written from, including the output of tools the policy names. The policy governs tool-call observations only.

The label model also receives that evidence. A host that sends labels to a different provider than its agents and relied on this policy to keep tool output away from that provider needs a different control.

buildActivityLabelPrompt, buildActivityPhaseLabelPrompt and buildReasoningLabelPrompt keep their optional redaction parameter for hosts that call them directly.

Tests

  • activity-label-redaction-scope.test.ts (new): under a run-level run_select_query policy, an activity label receives the tool output, reasoning and previous headers; a reasoning-only block is labeled; a labels-only phase is titled from its child labels and commentary; a reasoning label is produced. All four fail against main.
  • activity-phase-label.test.ts and reasoning-label.test.ts: three cases that asserted the label model saw redacted or no evidence now assert it sees the full evidence.

The full suite passes apart from three provider suites that need live credentials (llm/anthropic, llm/google, llm/vertexai): 278 suites, 5973 tests. tsc --noEmit and ESLint are clean on the changed files.

Built into a LibreChat checkout that sets a run_select_query / run_tools_with_bash redaction policy with phase labels enabled, this turns the activity-fold.spec.ts e2e suite from 2 failed to 2 passed with the policy still in place.

Supersedes #573.

@danny-avila
danny-avila merged commit e269670 into main Sep 28, 2026
@danny-avila danny-avila changed the title 🙈 fix: Redact Label Evidence in Traces, Not in the Label Prompt 🏷️ fix: Scope Tool-Output Redaction to Tool Calls, Not Label Evidence Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant