Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ base:
name: qwen3.5-fp8-mi355x-sglang-mtp-8k1k
model:
path: hf:Qwen/Qwen3.5-397B-A17B-FP8
container: lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260828
container: lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260924
precision: fp8
resources:
gpu_type: mi355x
Expand All @@ -29,7 +29,7 @@ base:
tokenizer-worker-num: 6
enable-aiter-allreduce-fusion: true
max-running-requests: 4
cuda-graph-max-bs: 4
cuda-graph-max-bs-decode: 4
disable-radix-cache: true
chunked-prefill-size: 32768
scheduler-recv-interval: 30
Expand Down Expand Up @@ -60,7 +60,7 @@ zip_override_concurrency:
roles:
agg:
args:
cuda-graph-max-bs: [4, 8, 16, 32, 64, 128, 256]
cuda-graph-max-bs-decode: [4, 8, 16, 32, 64, 128, 256]
max-running-requests: [4, 8, 16, 32, 64, 128, 256]
benchmark:
env:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ base:
name: qwen3.5-fp8-mi355x-sglang-8k1k
model:
path: hf:Qwen/Qwen3.5-397B-A17B-FP8
container: lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260828
container: lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260924
precision: fp8
resources:
gpu_type: mi355x
Expand All @@ -29,7 +29,7 @@ base:
tokenizer-worker-num: 6
enable-aiter-allreduce-fusion: true
max-running-requests: 4
cuda-graph-max-bs: 4
cuda-graph-max-bs-decode: 4
disable-radix-cache: true
chunked-prefill-size: 32768
scheduler-recv-interval: 30
Expand All @@ -56,7 +56,7 @@ zip_override_concurrency:
roles:
agg:
args:
cuda-graph-max-bs: [4, 8, 16, 32, 64, 128, 256]
cuda-graph-max-bs-decode: [4, 8, 16, 32, 64, 128, 256]
max-running-requests: [4, 8, 16, 32, 64, 128, 256]
benchmark:
env:
Expand Down
4 changes: 2 additions & 2 deletions configs/amd-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -182,7 +182,7 @@ qwen3.5-fp8-mi325x-sglang:
srt-recipe: benchmarks/single_node/srt-slurm-recipes/qwen3.5/sglang/mi325x-fp8/8k1k.yaml

qwen3.5-fp8-mi355x-sglang:
image: lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260828
image: lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260924
model: Qwen/Qwen3.5-397B-A17B-FP8
model-prefix: qwen3.5
runner: mi355x
Expand All @@ -201,7 +201,7 @@ qwen3.5-fp8-mi355x-sglang:
srt-recipe: benchmarks/single_node/srt-slurm-recipes/qwen3.5/sglang/mi355x-fp8/8k1k.yaml

qwen3.5-fp8-mi355x-sglang-mtp:
image: lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260828
image: lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260924
model: Qwen/Qwen3.5-397B-A17B-FP8
model-prefix: qwen3.5
runner: mi355x
Expand Down
7 changes: 7 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8902,3 +8902,10 @@
- "server_atom.sh exports PORT before sourcing benchmark_lib.sh on the eval path. run_lm_eval's check_env_vars guard runs before it parses --port and job.slurm's docker -e allowlist forwards ROUTER_PORT but never PORT, so every multi-node lm-eval cell aborted 8 s in; because check_env_vars exits rather than returning, node 0 also skipped its router teardown and the decode node hung in 'Waiting until router closes...' until the job was cancelled (run 35643358891)."
- "Per ATOM DeepSeek-V4-Agentic-PD-Max recipe: prefix caching on, FP8 KV and index cache, block-size 256, TBO on prefill only, max-num-seqs = 2x concurrency."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3158

- config-keys:
- qwen3.5-fp8-mi355x-sglang
- qwen3.5-fp8-mi355x-sglang-mtp
description:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The new entry's pr-link: PRLINK_PLACEHOLDER will make the merge-prep automation reject this PR instead of auto-filling the real link. infx/workflows/prepare_perf_changelog_merge.py (validate_perf_changelog.py) only accepts "XXX" or "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX" as placeholders (PR_LINK_PLACEHOLDERS); any other value raises ChangelogValidationError(f"appended entry {index + 1} has unexpected pr-link {link!r}"). Fix: use the documented placeholder https://github.com/SemiAnalysisAI/InferenceX/pull/XXX (or XXX) so the merge automation can canonicalize it to the real PR link.

Why this was flagged

perf-changelog.yaml:8894 sets pr-link: PRLINK_PLACEHOLDER for the newly appended entry (config-keys qwen3.5-fp8-mi355x-sglang / -mtp). infx/workflows/prepare_perf_changelog_merge.py:96-109 computes expected_link and, for each appended entry whose link isn't already the expected link, requires it be in PR_LINK_PLACEHOLDERS = {"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"} (validate_perf_changelog.py:21-24); otherwise it raises ChangelogValidationError. PRLINK_PLACEHOLDER matches neither, so the merge-prep step (and validate_perf_changelog.py:138's same check) fails, blocking this PR from merging via the normal automated path where the base branch would have a correctly formatted placeholder or real link.

Verification: normal. The appended entry sets pr-link: PRLINK_PLACEHOLDER (perf-changelog.yaml:8896). The merge-prep automation infx/workflows/prepare_perf_changelog_merge.py command canonicalize calls canonicalize_appended_links -> compare_entries(base, head, pr_number) (line 86). Inside compare_entries, every appended entry passes through validate_added_pr_link(link, pr_number)… | normal. The…

- "Update SGLang ROCm image from v0.5.18-rocm720-mi35x-20260828 to v0.5.20-rocm720-mi35x-20260924 (latest nightly)"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3423
Loading