Skip to content

Commit 8326784

Browse files
authored
feat(api): Add prompt cache diagnostics (openai#3800)
## Summary Expose optional prompt cache comparison diagnostics across the stable and beta Responses APIs. Clarify how Fast and Priority service tier requests resolve for different models. ## Changes - Add the optional `comparison_response_id` prompt cache option to response creation parameters and client events. - Add `prompt_cache_diagnostics` to response models with typed cache hit, cache miss, comparison-not-found, and unavailable outcomes. - Expose cache miss reasons, affected token estimates, and reusable-prefix token counts in diagnostic results. - Include the comparison response ID in returned prompt cache options when supplied. - Clarify that Fast or Priority requests resolve to `fast` for models with a dedicated Fast tier and to `priority` for other models. Co-authored-by: apcha-oai <228803254+apcha-oai@users.noreply.github.com>
1 parent 3cc8d78 commit 8326784

16 files changed

Lines changed: 517 additions & 55 deletions

‎.castiron.stats.yml‎

Lines changed: 7 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,8 @@
11
schema_version: 1
2-
generation_id: 08cffa97-63b1-4ab3-b3e4-9d438f45458f
3-
openapi_spec_hash: fcdb9c99a509f5ea8b90d7abbf1dab2f
4-
openapi_transformed_spec_hash: dd244f7dd3ab9d18dc2c34970a39505f
5-
config_hash: 81fd1039eb373628d87269f065d30a69
6-
codegen_sha: fb6674883ea2eed44fa6db1f16b106ee84b4c9e5
7-
codegen_hash: d68cd064823d9541b5d0874c92142d2691f7943452d9142c205b4cead3798785
8-
public_codegen_sha: e84368fa07e999995b4ec9f8514e6c5a98be2a89
2+
generation_id: 50d48205-978b-4b96-83b9-e92c82968d25
3+
openapi_spec_hash: a5a64d9aaf4a4022898dd86685c102ab
4+
openapi_transformed_spec_hash: f7c7c9ee2566bcf8165ab87b4851a1e5
5+
config_hash: dfe8a20c64cc5d852e69c441cbd28ae7
6+
codegen_sha: 9cbf6087efd2a6d5a57028c17f031fea6f7f5015
7+
codegen_hash: 331343ebbad785fc4c6d5ce8197e5b3de60b4b989f77df08356eb00fc6e958ef
8+
public_codegen_sha: 0c33f635083e108a823a23805ec04f2e6d1add4d

‎api_reference/openapi.transformed.yml‎

Lines changed: 298 additions & 12 deletions
Large diffs are not rendered by default.

‎src/openai/resources/beta/responses/responses.py‎

Lines changed: 14 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -1958,12 +1958,13 @@ def compact(
19581958
request will be processed with the Flex Processing service tier. - To opt-in to
19591959
[Fast mode](/api/docs/guides/fast-mode) at the request level, include the
19601960
`service_tier=fast` or `service_tier=priority` parameter for Responses or Chat
1961-
Completions. The response will show `service_tier=priority` regardless of if you
1962-
specify `service_tier=fast` or `priority` in your request. - When not set, the
1963-
default behavior is 'auto'. When the `service_tier` parameter is set, the
1964-
response body will include the `service_tier` value based on the processing mode
1965-
actually used to serve the request. This response value may be different from
1966-
the value set in the parameter.
1961+
Completions. For models with a dedicated Fast tier, either value resolves to
1962+
`service_tier=fast`; for other models, either value resolves to
1963+
`service_tier=priority`. - When not set, the default behavior is 'auto'. When
1964+
the `service_tier` parameter is set, the response body will include the
1965+
`service_tier` value based on the processing mode actually used to serve the
1966+
request. This response value may be different from the value set in the
1967+
parameter.
19671968
19681969
extra_headers: Send extra headers
19691970
@@ -3912,12 +3913,13 @@ async def compact(
39123913
request will be processed with the Flex Processing service tier. - To opt-in to
39133914
[Fast mode](/api/docs/guides/fast-mode) at the request level, include the
39143915
`service_tier=fast` or `service_tier=priority` parameter for Responses or Chat
3915-
Completions. The response will show `service_tier=priority` regardless of if you
3916-
specify `service_tier=fast` or `priority` in your request. - When not set, the
3917-
default behavior is 'auto'. When the `service_tier` parameter is set, the
3918-
response body will include the `service_tier` value based on the processing mode
3919-
actually used to serve the request. This response value may be different from
3920-
the value set in the parameter.
3916+
Completions. For models with a dedicated Fast tier, either value resolves to
3917+
`service_tier=fast`; for other models, either value resolves to
3918+
`service_tier=priority`. - When not set, the default behavior is 'auto'. When
3919+
the `service_tier` parameter is set, the response body will include the
3920+
`service_tier` value based on the processing mode actually used to serve the
3921+
request. This response value may be different from the value set in the
3922+
parameter.
39213923
39223924
extra_headers: Send extra headers
39233925

‎src/openai/resources/responses/responses.py‎

Lines changed: 14 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -1897,12 +1897,13 @@ def compact(
18971897
request will be processed with the Flex Processing service tier. - To opt-in to
18981898
[Fast mode](/api/docs/guides/fast-mode) at the request level, include the
18991899
`service_tier=fast` or `service_tier=priority` parameter for Responses or Chat
1900-
Completions. The response will show `service_tier=priority` regardless of if you
1901-
specify `service_tier=fast` or `priority` in your request. - When not set, the
1902-
default behavior is 'auto'. When the `service_tier` parameter is set, the
1903-
response body will include the `service_tier` value based on the processing mode
1904-
actually used to serve the request. This response value may be different from
1905-
the value set in the parameter.
1900+
Completions. For models with a dedicated Fast tier, either value resolves to
1901+
`service_tier=fast`; for other models, either value resolves to
1902+
`service_tier=priority`. - When not set, the default behavior is 'auto'. When
1903+
the `service_tier` parameter is set, the response body will include the
1904+
`service_tier` value based on the processing mode actually used to serve the
1905+
request. This response value may be different from the value set in the
1906+
parameter.
19061907
19071908
extra_headers: Send extra headers
19081909
@@ -3759,12 +3760,13 @@ async def compact(
37593760
request will be processed with the Flex Processing service tier. - To opt-in to
37603761
[Fast mode](/api/docs/guides/fast-mode) at the request level, include the
37613762
`service_tier=fast` or `service_tier=priority` parameter for Responses or Chat
3762-
Completions. The response will show `service_tier=priority` regardless of if you
3763-
specify `service_tier=fast` or `priority` in your request. - When not set, the
3764-
default behavior is 'auto'. When the `service_tier` parameter is set, the
3765-
response body will include the `service_tier` value based on the processing mode
3766-
actually used to serve the request. This response value may be different from
3767-
the value set in the parameter.
3763+
Completions. For models with a dedicated Fast tier, either value resolves to
3764+
`service_tier=fast`; for other models, either value resolves to
3765+
`service_tier=priority`. - When not set, the default behavior is 'auto'. When
3766+
the `service_tier` parameter is set, the response body will include the
3767+
`service_tier` value based on the processing mode actually used to serve the
3768+
request. This response value may be different from the value set in the
3769+
parameter.
37683770
37693771
extra_headers: Send extra headers
37703772

‎src/openai/types/beta/beta_response.py‎

Lines changed: 60 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -36,6 +36,11 @@
3636
"ModerationOutput",
3737
"ModerationOutputModerationResult",
3838
"ModerationOutputError",
39+
"PromptCacheDiagnostics",
40+
"PromptCacheDiagnosticsCacheMiss",
41+
"PromptCacheDiagnosticsCacheHit",
42+
"PromptCacheDiagnosticsComparisonResponseNotFound",
43+
"PromptCacheDiagnosticsUnavailable",
3944
"PromptCacheOptions",
4045
"Reasoning",
4146
]
@@ -185,6 +190,55 @@ class Moderation(BaseModel):
185190
"""Moderation for the response output."""
186191

187192

193+
class PromptCacheDiagnosticsCacheMiss(BaseModel):
194+
cache_missed_tokens: int
195+
"""
196+
The estimated number of input tokens affected after the first detected
197+
divergence.
198+
"""
199+
200+
reason: Literal[
201+
"model_changed",
202+
"prompt_cache_key_changed",
203+
"tools_changed",
204+
"text_format_changed",
205+
"reasoning_effort_changed",
206+
"verbosity_changed",
207+
"context_compacted",
208+
"input_changed",
209+
"service_tier_changed",
210+
]
211+
"""The reason prompt cache reuse did not occur."""
212+
213+
type: Literal["cache_miss"]
214+
215+
comparison_reusable_tokens: Optional[int] = None
216+
"""The raw token count of the reusable prefix in the compared response."""
217+
218+
219+
class PromptCacheDiagnosticsCacheHit(BaseModel):
220+
type: Literal["cache_hit"]
221+
222+
223+
class PromptCacheDiagnosticsComparisonResponseNotFound(BaseModel):
224+
type: Literal["comparison_response_not_found"]
225+
226+
227+
class PromptCacheDiagnosticsUnavailable(BaseModel):
228+
type: Literal["unavailable"]
229+
230+
231+
PromptCacheDiagnostics: TypeAlias = Annotated[
232+
Union[
233+
PromptCacheDiagnosticsCacheMiss,
234+
PromptCacheDiagnosticsCacheHit,
235+
PromptCacheDiagnosticsComparisonResponseNotFound,
236+
PromptCacheDiagnosticsUnavailable,
237+
],
238+
PropertyInfo(discriminator="type"),
239+
]
240+
241+
188242
class PromptCacheOptions(BaseModel):
189243
"""The prompt-caching options that were applied to the response.
190244
@@ -197,6 +251,9 @@ class PromptCacheOptions(BaseModel):
197251
ttl: Literal["30m"]
198252
"""The minimum lifetime applied to each cache breakpoint."""
199253

254+
comparison_response_id: Optional[str] = None
255+
"""The response ID supplied as the prompt cache diagnostics comparison."""
256+
200257

201258
class Reasoning(BaseModel):
202259
"""
@@ -514,6 +571,9 @@ class BetaResponse(BaseModel):
514571
[Learn more](https://platform.openai.com/docs/guides/text?api-mode=responses#reusable-prompts).
515572
"""
516573

574+
prompt_cache_diagnostics: Optional[PromptCacheDiagnostics] = None
575+
"""Prompt cache diagnostics requested for this response."""
576+
517577
prompt_cache_key: Optional[str] = None
518578
"""
519579
Used by OpenAI to cache responses for similar requests to optimize your cache

‎src/openai/types/beta/beta_responses_client_event.py‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -110,6 +110,13 @@ class ResponseCreatePromptCacheOptions(BaseModel):
110110
Supported for `gpt-5.6` and later models. By default, OpenAI automatically chooses one implicit cache breakpoint. You can add explicit breakpoints to content blocks with `prompt_cache_breakpoint`. Each request can write up to four breakpoints. For cache matching, OpenAI considers up to the latest 80 breakpoints in the conversation, without a content-block lookback limit. Set `mode` to `explicit` to disable the implicit breakpoint. The `ttl` defaults to `30m`, which is currently the only supported value. See the [prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching) for current details.
111111
"""
112112

113+
comparison_response_id: Optional[str] = None
114+
"""The ID of a response to compare when diagnosing prompt cache reuse.
115+
116+
Supplying this field requests prompt cache diagnostics when the feature is
117+
enabled.
118+
"""
119+
113120
mode: Optional[Literal["implicit", "explicit"]] = None
114121
"""Controls whether OpenAI automatically creates an implicit cache breakpoint.
115122

‎src/openai/types/beta/beta_responses_client_event_param.py‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -110,6 +110,13 @@ class ResponseCreatePromptCacheOptions(TypedDict, total=False):
110110
Supported for `gpt-5.6` and later models. By default, OpenAI automatically chooses one implicit cache breakpoint. You can add explicit breakpoints to content blocks with `prompt_cache_breakpoint`. Each request can write up to four breakpoints. For cache matching, OpenAI considers up to the latest 80 breakpoints in the conversation, without a content-block lookback limit. Set `mode` to `explicit` to disable the implicit breakpoint. The `ttl` defaults to `30m`, which is currently the only supported value. See the [prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching) for current details.
111111
"""
112112

113+
comparison_response_id: Optional[str]
114+
"""The ID of a response to compare when diagnosing prompt cache reuse.
115+
116+
Supplying this field requests prompt cache diagnostics when the feature is
117+
enabled.
118+
"""
119+
113120
mode: Literal["implicit", "explicit"]
114121
"""Controls whether OpenAI automatically creates an implicit cache breakpoint.
115122

‎src/openai/types/beta/response_compact_params.py‎

Lines changed: 7 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -181,12 +181,13 @@ class ResponseCompactParams(TypedDict, total=False):
181181
request will be processed with the Flex Processing service tier. - To opt-in
182182
to [Fast mode](/api/docs/guides/fast-mode) at the request level, include the
183183
`service_tier=fast` or `service_tier=priority` parameter for Responses or Chat
184-
Completions. The response will show `service_tier=priority` regardless of if
185-
you specify `service_tier=fast` or `priority` in your request. - When not set,
186-
the default behavior is 'auto'. When the `service_tier` parameter is set, the
187-
response body will include the `service_tier` value based on the processing
188-
mode actually used to serve the request. This response value may be different
189-
from the value set in the parameter.
184+
Completions. For models with a dedicated Fast tier, either value resolves to
185+
`service_tier=fast`; for other models, either value resolves to
186+
`service_tier=priority`. - When not set, the default behavior is 'auto'. When
187+
the `service_tier` parameter is set, the response body will include the
188+
`service_tier` value based on the processing mode actually used to serve the
189+
request. This response value may be different from the value set in the
190+
parameter.
190191
"""
191192

192193
betas: Annotated[List[Literal["responses_multi_agent=v1"]], PropertyInfo(alias="openai-beta")]

‎src/openai/types/beta/response_create_params.py‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -509,6 +509,13 @@ class PromptCacheOptions(TypedDict, total=False):
509509
Supported for `gpt-5.6` and later models. By default, OpenAI automatically chooses one implicit cache breakpoint. You can add explicit breakpoints to content blocks with `prompt_cache_breakpoint`. Each request can write up to four breakpoints. For cache matching, OpenAI considers up to the latest 80 breakpoints in the conversation, without a content-block lookback limit. Set `mode` to `explicit` to disable the implicit breakpoint. The `ttl` defaults to `30m`, which is currently the only supported value. See the [prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching) for current details.
510510
"""
511511

512+
comparison_response_id: Optional[str]
513+
"""The ID of a response to compare when diagnosing prompt cache reuse.
514+
515+
Supplying this field requests prompt cache diagnostics when the feature is
516+
enabled.
517+
"""
518+
512519
mode: Literal["implicit", "explicit"]
513520
"""Controls whether OpenAI automatically creates an implicit cache breakpoint.
514521

‎src/openai/types/responses/response.py‎

Lines changed: 60 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -39,6 +39,11 @@
3939
"ModerationOutput",
4040
"ModerationOutputModerationResult",
4141
"ModerationOutputError",
42+
"PromptCacheDiagnostics",
43+
"PromptCacheDiagnosticsCacheMiss",
44+
"PromptCacheDiagnosticsCacheHit",
45+
"PromptCacheDiagnosticsComparisonResponseNotFound",
46+
"PromptCacheDiagnosticsUnavailable",
4247
"PromptCacheOptions",
4348
]
4449

@@ -187,6 +192,55 @@ class Moderation(BaseModel):
187192
"""Moderation for the response output."""
188193

189194

195+
class PromptCacheDiagnosticsCacheMiss(BaseModel):
196+
cache_missed_tokens: int
197+
"""
198+
The estimated number of input tokens affected after the first detected
199+
divergence.
200+
"""
201+
202+
reason: Literal[
203+
"model_changed",
204+
"prompt_cache_key_changed",
205+
"tools_changed",
206+
"text_format_changed",
207+
"reasoning_effort_changed",
208+
"verbosity_changed",
209+
"context_compacted",
210+
"input_changed",
211+
"service_tier_changed",
212+
]
213+
"""The reason prompt cache reuse did not occur."""
214+
215+
type: Literal["cache_miss"]
216+
217+
comparison_reusable_tokens: Optional[int] = None
218+
"""The raw token count of the reusable prefix in the compared response."""
219+
220+
221+
class PromptCacheDiagnosticsCacheHit(BaseModel):
222+
type: Literal["cache_hit"]
223+
224+
225+
class PromptCacheDiagnosticsComparisonResponseNotFound(BaseModel):
226+
type: Literal["comparison_response_not_found"]
227+
228+
229+
class PromptCacheDiagnosticsUnavailable(BaseModel):
230+
type: Literal["unavailable"]
231+
232+
233+
PromptCacheDiagnostics: TypeAlias = Annotated[
234+
Union[
235+
PromptCacheDiagnosticsCacheMiss,
236+
PromptCacheDiagnosticsCacheHit,
237+
PromptCacheDiagnosticsComparisonResponseNotFound,
238+
PromptCacheDiagnosticsUnavailable,
239+
],
240+
PropertyInfo(discriminator="type"),
241+
]
242+
243+
190244
class PromptCacheOptions(BaseModel):
191245
"""The prompt-caching options that were applied to the response.
192246
@@ -199,6 +253,9 @@ class PromptCacheOptions(BaseModel):
199253
ttl: Literal["30m"]
200254
"""The minimum lifetime applied to each cache breakpoint."""
201255

256+
comparison_response_id: Optional[str] = None
257+
"""The response ID supplied as the prompt cache diagnostics comparison."""
258+
202259

203260
class Response(BaseModel):
204261
id: str
@@ -357,6 +414,9 @@ class Response(BaseModel):
357414
[Learn more](https://platform.openai.com/docs/guides/text?api-mode=responses#reusable-prompts).
358415
"""
359416

417+
prompt_cache_diagnostics: Optional[PromptCacheDiagnostics] = None
418+
"""Prompt cache diagnostics requested for this response."""
419+
360420
prompt_cache_key: Optional[str] = None
361421
"""
362422
Used by OpenAI to cache responses for similar requests to optimize your cache

0 commit comments

Comments
 (0)