Conversation
…t omit them ChatCompletionStreamState._accumulate_chunk assigned both from every chunk unconditionally, so a trailing chunk that omits them wiped values the stream had already reported from the snapshot and from get_final_completion(). Guard both the way the adjacent moderation assignment already does.
|
Independent offline verification (AI-assisted): I independently exercised the actual Compared main Reproducer, runnable separately against each source tree: import json
from openai.types.chat import ChatCompletionChunk
from openai.lib.streaming.chat import ChatCompletionStreamState
import openai.lib.streaming.chat._completions as source
sequences=[('early_then_omitted',[(11,'fp-a'),(None,None),(None,None)]),('last_only',[(None,None),(None,None),(11,'fp-a')]),('later_update',[(11,'fp-a'),(13,'fp-b'),(None,None)]),('usage_only',[(11,None),(None,None)]),('fingerprint_only',[(None,'fp-a'),(None,None)]),('never_reported',[(None,None),(None,None)]),('zero_usage',[(0,'fp-a'),(None,None)])]
rows=[]
for label,parts in sequences:
state=ChatCompletionStreamState();expected_usage=None;expected_fp=None;observations=[]
for i,(tokens,fp) in enumerate(parts):
usage=None if tokens is None else {'prompt_tokens':tokens,'completion_tokens':0,'total_tokens':tokens}
delta={'role':'assistant','content':'fictional'} if i==0 else {}
chunk=ChatCompletionChunk(id='fictional',object='chat.completion.chunk',created=0,model='fictional',choices=[{'index':0,'delta':delta,'finish_reason':'stop' if i==len(parts)-1 else None}],usage=usage,system_fingerprint=fp)
list(state.handle_chunk(chunk))
if tokens is not None:expected_usage=tokens
if fp is not None:expected_fp=fp
snapshot=state.current_completion_snapshot
observations.append({'actual_usage':snapshot.usage.total_tokens if snapshot.usage else None,'expected_usage':expected_usage,'actual_fp':snapshot.system_fingerprint,'expected_fp':expected_fp})
final=state.get_final_completion()
ok=all(o['actual_usage']==o['expected_usage'] and o['actual_fp']==o['expected_fp'] for o in observations)
ok=ok and (final.usage.total_tokens if final.usage else None)==expected_usage and final.system_fingerprint==expected_fp
rows.append({'case':label,'pass':ok,'snapshots':observations})
print(json.dumps({'source_module':source.__file__,'passed':sum(r['pass'] for r in rows),'failed':sum(not r['pass'] for r in rows),'cases':rows,'scope':'Actual ChatCompletionStreamState with fictional constructed chunks; not client SSE/provider proof. Zero totals are a constructed boundary control.'})) |
Changes being requested
ChatCompletionStreamState._accumulate_chunkwrotesnapshot.usageandsnapshot.system_fingerprintfrom every chunk unconditionally:The two lines above the
moderationguard overwrite with whatever the current chunk carries, which for most chunks isNone. So a stream that reports usage on one chunk and then sends a trailing chunk that omits it loses both fields — fromcurrent_completion_snapshot, from every subsequentChunkEvent.snapshot, and fromget_final_completion().That trailing-chunk shape is not hypothetical: it is what
tests/lib/chat/test_stream_moderation.pymodels, where a metadata chunk arrives after the content chunks. #3864 established the rule formoderation— a later chunk that omits the field retains the last report — but the two siblings immediately above it were left unconditional. This applies the same guard to both.src/openai/lib/streaming/chat/_completions.py+4/-2, plus a new testtests/lib/chat/test_stream_metadata_retention.pythat mirrors the moderation test's structure with usage and the fingerprint reported on the first chunk and a trailing moderation chunk after them.Reproduction
Before the change, driving the real
.stream()path through a mocked SSE response (usageandsystem_fingerprinton chunk 0, a trailing moderation chunk at index 2), the per-chunk snapshots ofusage.total_tokensare:Driving
ChatCompletionStreamStatedirectly with the same chunk order — which is what.stream()accumulates through — shows the final object losing both fields while the sibling keeps its value:The same stream with
usageon the last chunk loses nothing, which is why the existing suite did not catch it — the loss is order-dependent. Reported values are never cleared by this change: a later chunk that does report usage still replaces the earlier one.Verification
main@bccad312, openai 3.16.1, Python 3.10.16, pydantic 2.12.5, macOS arm64,uv sync --frozen --all-extrasthenuv run --frozen --no-sync. The new test file is present in both runs, so the totals are directly comparable.pytest -o addopts= -q)main, fix reverted)tests/lib/chat/test_stream_metadata_retention.pyAt index 1 diff: None != 11tests/lib/chat/test_stream_moderation.pytests/lib/chat tests/lib/streaming tests/lib/responsestests/lib(full)I diffed the sorted
FAILEDlists from the two full runs: the only difference is the two new tests disappearing from it. The remaining 40 failures are pre-existing and identical on both trees — all intests/lib/test_fine_tuning_positional_arguments.py, which needs a local echo server.ruff format --checkandruff checkare clean on both changed files.Notes
_accumulate_chunkis already guarded:finish_reason(if choice.finish_reason:),logprobs(if choice.logprobs is not None:) andmoderation(if chunk.moderation is not None:, added by fix: preserve chat stream moderation results #3864).usageandsystem_fingerprintwere the only two unconditional ones left.finish_reason == "length"the code raisesLengthFinishReasonError(completion=completion_snapshot)with the comment "at the time of writing,.usagewill always beNonebut we include it here in case that is changed in the future". After this change the snapshot can carry a retainedusage, so that error now reports accurate usage — which is what the comment anticipates, not a behaviour it depends on being absent.stream_options.include_usagethe API sendsusageon the final chunk, so the guard is a no-op there.src/openai/lib/; the Castiron report should show no new custom-code files and no changed customizations.moderationguard that fix: preserve chat stream moderation results #3864 already merged, so it does not cover these two fields.Additional context & links
Follows the rule #3864 set for
moderationin the same function.