| title | Validation Log |
|---|---|
| slug | validation-log |
| type | reference |
| status | canonical |
| created | 2026-04-21 |
| updated | 2026-05-07 |
| owner | Team 13 |
| scope | team-repo |
| canonical | true |
Last updated: 2026-05-07
Canonical log for live serve / benchmark / profiling proofs. Use this file for concrete run records, not the runbooks.
For each proof entry, record:
- date
- scope
- branch / git SHA
- config or command path
- run id / Slurm job id
- primary artifacts
- what the run proves
- caveats / follow-ups
- Scope: Hosted WatsonX Llama-3.3-70B scale-sanity runs for the post-PR175
clean-floor evidence set. This closes the exploratory 70B decision for
#95: 70B is appendix / scale-sanity evidence, not part of the core 8B matrix claim. - Model:
watsonx/meta-llama/llama-3-3-70b-instruct. - Branch / git SHA:
team13/main@1913c6e4703425f735d8cb8297cb890ba66bbeff. - Cohorts:
watsonx70b_main15pluswatsonx70b_topup15inresults/metrics/evidence_registry.csv. - Row groups:
A70B,B70B,C70B,Y70B,YS70B,Z70B,ZS70B. - Coverage: 11 registry rows; 435 raw trajectories; 435 judge rows. A/B/C have 15 scenarios x 3 trials; Y/YS/Z/ZS have 15 scenarios x 5 trials after topups.
- Primary artifacts:
results/metrics/evidence_registry.csvresults/metrics/gcp_post175_70b_summary.csvresults/metrics/scenario_scores.jsonlbenchmarks/cell_A70B/raw/final15x3_70b_watsonx_post175_cpu_ixqt_abc_20260505T0722Z_A70B_post175_70b_aat_direct_watsonxbenchmarks/cell_B70B/raw/final15x3_70b_watsonx_post175_cpu_ixqt_abc_20260505T0722Z_B70B_post175_70b_aat_mcp_baseline_watsonxbenchmarks/cell_C70B/raw/final15x3_70b_watsonx_post175_cpu_ixqt_abc_20260505T0722Z_C70B_post175_70b_aat_mcp_optimized_watsonxbenchmarks/cell_Y70B/raw/final15x3_70b_watsonx_post175_cpu_cu_a_yys_envrepair_20260505T0820Z_Y70B_post175_70b_exp2_cell_Y_pe_mcp_baseline_watsonxbenchmarks/cell_YS70B/raw/final15x3_70b_watsonx_post175_cpu_cu_a_yys_envrepair_20260505T0820Z_YS70B_post175_70b_exp2_cell_YS_pe_self_ask_mcp_baseline_watsonxbenchmarks/cell_Z70B/raw/final15x3_70b_watsonx_post175_cpu_cu_b_zzs_envrepair_20260505T0820Z_Z70B_post175_70b_exp2_cell_Z_verified_pe_mcp_baseline_watsonxbenchmarks/cell_ZS70B/raw/final15x3_70b_watsonx_post175_cpu_cu_b_zzs_envrepair_20260505T0820Z_ZS70B_post175_70b_exp2_cell_ZS_verified_pe_self_ask_mcp_baseline_watsonx
- Aggregate judge summary: 218 / 435 pass at threshold 0.6. Per-row pass/mean score: A70B 23/45, mean 0.5444; B70B 23/45, mean 0.5704; C70B 22/45, mean 0.5815; Y70B 37/75, mean 0.6244; YS70B 34/75, mean 0.5867; Z70B 40/75, mean 0.6200; ZS70B 39/75, mean 0.6444.
What this proves:
- The hosted 70B path can run and judge the same Smart Grid task surface under
the post-PR175 clean-floor code, with accepted rows registered as
paper_grade. - The scale-sanity trend does not overturn the 8B core story: 70B coverage is narrower than the core all-scenario matrix and is best used as an appendix cross-check rather than as a primary matrix claim.
Caveats / follow-ups:
- The core NeurIPS matrix remains the registry-backed 8B evidence set.
- Full all-scenario / all-generated-scenario 70B expansion remains future work
and is not a blocker for
#95.
- Scope: Exploratory best-engineered PE-family ablation: Verified PE + Self-Ask with optimized MCP persistent sessions and the Cell D compressed-INT8 / BF16 / fp8-KV / prefix-cache serving profile.
- Scenario set:
data/scenarios/multi_*.json, 2 scenarios × 3 trials = 6 trial artifacts. - Model: self-hosted
openai/Llama-3.1-8B-Instruct-int8through local vLLM on Insomnia. - Branch / git SHA:
team13/main@eb7019b3ebeb3e3b6848ab4b0c8b4bf0fa21ab44in the shared Insomnia checkout. - Config:
configs/experiment2/exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized.env. - Run id / Slurm job id:
9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized. - Node / Slurm state:
ins084,COMPLETED 0:0, elapsed00:12:38. - W&B: https://wandb.ai/assetopsbench-smartgrid/assetopsbench-smartgrid/runs/48nqpclw
- Primary artifacts:
benchmarks/cell_ZSD/raw/9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized/meta.jsonbenchmarks/cell_ZSD/raw/9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized/harness.logbenchmarks/cell_ZSD/raw/9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized/vllm.logbenchmarks/cell_ZSD/raw/9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized/latencies.jsonlbenchmarks/cell_ZSD/config.json,benchmarks/cell_ZSD/summary.jsonresults/metrics/scenario_scores.jsonlresults/judge_logs/9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized/SGT-009_run{01,02,03}_judge_log.jsonresults/judge_logs/9074775_exp2_cell_ZSD_verified_pe_self_ask_mcp_model_optimized/SGT-010_run{01,02,03}_judge_log.json- W&B profiling artifact
profiling-48nqpclw
- Per-trial latencies:
114.28,78.44,86.72,31.89,30.41,29.47s. - Judge result: Maverick-17B six-dimension judge scored all six ZSD trials
with
score_6dvalues0.3333,0.5,0.0,0.8333,1.0,1.0; mean0.6111, p500.8333, pass rate3 / 6at threshold0.6.
What this proves:
- The PE-family optimized MCP extension is real on Insomnia: the runner reused
initialized MCP sessions, reached all four Smart Grid MCP servers, executed
model/tool loops, and completed
6 / 6withtool_error_count=0. - The model-side stack was the intended Cell D serving profile.
vllm.logrecordsdtype=torch.bfloat16,quantization=compressed-tensors,kv_cache_dtype=fp8,enable_prefix_caching=True, and the compressed-tensor Cutlass INT8 kernel. - The ZSD run is slower than the AaT optimized-serving arm but materially
better on the current six-dimension judge than Cells C/D: mean
0.6111and3 / 6judge-pass versus D mean0.1667and1 / 6.
Caveats / follow-ups:
- This is an exploratory ablation, not part of the clean core matrix. It stacks verifier logic, Self-Ask, optimized MCP persistent sessions, and model-side serving changes, so it is best interpreted as a PE-family ceiling.
- Failed boundary runs are intentionally not committed as proof artifacts:
9073604reached model/tool execution but had partial completion and stdout JSON pollution;9074217exposed an Insomnia AOB portability mismatch in the firstMAX_TOKENSwrapper. Commits9be831bandeb7019bfixed those boundaries before the successful9074775rerun. - Final paper-grade reruns should still freeze a scenario set and trial count across the core matrix before promoting this beyond an ablation/sidebar.
- Scope: Exploratory Cell D optimized-serving AaT capture for the
multi_*.jsonslice. Cell D uses the Cell C optimized MCP transport and additionally changes the serving stack to compressed INT8 weights, BF16 execution, fp8 KV cache, and prefix caching. - Scenario set:
data/scenarios/multi_*.json, 2 scenarios × 3 trials = 6 trial artifacts. - Model: self-hosted
openai/Llama-3.1-8B-Instruct-int8through local vLLM on Insomnia. - Branch / git SHA:
team13/main@ec17dc781f0d703fa365ba348e82ed4b38afb222in the shared Insomnia checkout for the run. - Config:
configs/aat_mcp_model_optimized.env. - Run id / Slurm job id:
9073472_aat_mcp_model_optimized. - Node / Slurm state:
ins084,COMPLETED 0:0, elapsed00:10:01. - W&B: https://wandb.ai/assetopsbench-smartgrid/assetopsbench-smartgrid/runs/pmwzatie
- Primary artifacts:
benchmarks/cell_D/raw/9073472_aat_mcp_model_optimized/meta.jsonbenchmarks/cell_D/raw/9073472_aat_mcp_model_optimized/harness.logbenchmarks/cell_D/raw/9073472_aat_mcp_model_optimized/vllm.logbenchmarks/cell_D/raw/9073472_aat_mcp_model_optimized/latencies.jsonlbenchmarks/cell_D/raw/9073472_aat_mcp_model_optimized/_batch_latencies.jsonlbenchmarks/cell_D/raw/9073472_aat_mcp_model_optimized/replay/replay_meta.jsonbenchmarks/cell_D/config.json,benchmarks/cell_D/summary.jsonresults/metrics/scenario_scores.jsonlresults/judge_logs/9073472_aat_mcp_model_optimized/SGT-009_run{01,02,03}_judge_log.jsonresults/judge_logs/9073472_aat_mcp_model_optimized/SGT-010_run{01,02,03}_judge_log.json- profiler trace directory
profiling/traces/9073472_aat_mcp_model_optimized_torch - W&B profiling artifact
profiling-pmwzatie
- Per-trial latencies:
18.67,8.05,6.30,3.77,6.04,4.77s with one-timemcp_setup_seconds=14.64. - Judge result: Maverick-17B six-dimension judge scored all six Cell D
trials with
score_6dvalues0.0,0.0,0.0,0.3333,0.6667,0.0; mean0.1667, p500.0, pass rate1 / 6at threshold0.6.
What this proves:
- Cell D runs end-to-end on the same first-capture Experiment 1 slice:
run_status: "success",6 / 6complete,tool_error_count=0, andtool_call_count_total=18. - The serving stack loaded the INT8 checkpoint and vLLM selected
CutlassInt8ScaledMMLinearKernel for CompressedTensorsW8A8Int8. Thevllm.logengine config recordsdtype=torch.bfloat16,quantization=compressed-tensors,kv_cache_dtype=fp8, andenable_prefix_caching=True. - Replay/profiling completed after the benchmark run:
replay_scenariospassed both unique scenarios and the torch profiler trace was linked to W&B. - Post-hoc LLM-as-judge scoring completed for all six per-trial trajectories and emitted per-trial audit logs.
Caveats / follow-ups:
- Cell D is not part of the clean A/B/C transport-only fairness contract. It is an optimized-serving ablation that changes model precision/quantization in addition to MCP transport.
- The run was emitted before commit
35efc6fexportedVLLM_DTYPE/EXTRA_VLLM_ARGSfor the Python metadata writer. The syncedbenchmarks/cell_D/config.jsonand per-runmeta.jsonhave an explicitmetadata_correction_note;harness.logandvllm.logare the primary proof that the serving stack used BF16, compressed tensors, fp8 KV cache, and prefix caching. Commit35efc6ffixes future runs, including ZSD. - Quality remains poor despite the faster optimized-serving execution. The one passing judge trial is not enough to promote D beyond exploratory status; use this as latency/systems evidence and motivation for PE-family optimized follow-ons.
- Scope: Experiment 1 Cell C optimized MCP capture for the
multi_*.jsonslice, matching the Cell B baseline depth from job8979314 - Scenario set:
data/scenarios/multi_*.json, 2 scenarios × 3 trials = 6 trial artifacts - Model: self-hosted
openai/Llama-3.1-8B-Instructthrough local vLLM on Insomnia - Branch / git SHA:
team13/main@7e8d169c7c0789378762b59219694e9b83964b67in the shared Insomnia checkout - Config:
configs/aat_mcp_optimized_singletool_runtime.envduring the proof run; this is the mergedconfigs/aat_mcp_optimized.envwithAAT_PARALLEL_TOOL_CALLS=falseafter the local vLLM path rejected parallel tool calls - Run id / Slurm job id:
9071639_aat_mcp_optimized - Node / Slurm state:
ins083,COMPLETED 0:0, elapsed00:19:07 - W&B: https://wandb.ai/assetopsbench-smartgrid/assetopsbench-smartgrid/runs/ifz8xfhm
- Primary artifacts: live artifacts in the shared Insomnia checkout:
benchmarks/cell_C_mcp_optimized/raw/9071639_aat_mcp_optimized/meta.jsonbenchmarks/cell_C_mcp_optimized/raw/9071639_aat_mcp_optimized/harness.logbenchmarks/cell_C_mcp_optimized/raw/9071639_aat_mcp_optimized/vllm.logbenchmarks/cell_C_mcp_optimized/raw/9071639_aat_mcp_optimized/latencies.jsonlbenchmarks/cell_C_mcp_optimized/raw/9071639_aat_mcp_optimized/_batch_latencies.jsonlbenchmarks/cell_C_mcp_optimized/raw/9071639_aat_mcp_optimized/replay/replay_meta.jsonbenchmarks/cell_C_mcp_optimized/config.json,benchmarks/cell_C_mcp_optimized/summary.jsonresults/metrics/scenario_scores.jsonlresults/judge_logs/9071639_aat_mcp_optimized/SGT-009_run{01,02,03}_judge_log.jsonresults/judge_logs/9071639_aat_mcp_optimized/SGT-010_run{01,02,03}_judge_log.json- profiler trace directory
profiling/traces/9071639_aat_mcp_optimized_torch - W&B profiling artifact
profiling-ifz8xfhm
- Per-trial latencies:
60.24,10.99,7.81,6.87,6.67,6.99s with one-timemcp_setup_seconds=39.07 - Comparison anchor: Cell B baseline job
8979314_aat_mcp_baseline(run_status: success, 6 / 6, p5012.91s, mean13.38s) - Judge result: Maverick-17B six-dimension judge scored all six Cell C
trials with
score_6dvalues0.1667,0.1667,0.0,0.1667,0.1667,0.3333; mean0.1667, p500.1667, pass rate0 / 6at threshold0.6
What this proves:
- Cell C now runs end-to-end on the canonical multi-scenario Experiment 1 slice:
run_status: "success",6 / 6complete,tool_error_count=0, andtool_call_count_total=18 - the optimized batch runner keeps the four Smart Grid MCP servers alive across
all six trials and writes canonical per-trial JSON,
_batch_latencies.jsonl,latencies.jsonl,summary.json, andmeta.json - vLLM prefix caching was active in the serving layer;
vllm.logreported prefix-cache hit rates rising through the run - replay/profiling completed after the benchmark run:
replay_scenariospassed both unique scenarios and the torch profiler trace was linked to W&B - post-hoc LLM-as-judge scoring completed for all six per-trial trajectories and emitted per-trial audit logs
- Notebook 02 can now compute the first real
(B - C)MCP-overhead headline. Using the aggregate summary fields, Cell C p50 was6.99s versus Cell B p5012.91s (about5.92s faster at p50). The Cell C mean was slower (16.60s versus13.38s) because the first optimized trial paid a large cold-start / first-prefix cost; excluding that first Cell C trial gives a steady-state mean near7.87s.
Caveats / follow-ups:
- job
9071621_aat_mcp_optimizedis negative evidence forAAT_PARALLEL_TOOL_CALLS=true: it reached vLLM, MCP bootstrap, model requests, and tool execution, but all six trials failed withThis model only supports single tool-calls at once! - the canonical Cell C config should keep
AAT_PARALLEL_TOOL_CALLS=falsefor the Insomnia vLLM / Llama-3.1-8B-Instruct path unless a future model/parser combination proves true parallel tool-call support ins082should stay excluded on future Insomnia submits; the earlier9071602attempt failed before vLLM because the node could not expose a GPU tonvidia-smi- the execution path is clean, but the judge result is poor:
0 / 6trajectories pass the currentscore_6d >= 0.6threshold, mainly because the agent failed to retrieve/ground the required evidence before answering - this is the first successful Cell C capture, not the final paper-grade 5-trial run set; final reruns should use the same scenario/model surface as the final A/B captures
- Scope: Agent-as-Tool smoke proofs for Experiment 1 Cell A (direct Python
tools), Cell B (MCP baseline), and the upstream
OpenAIAgentRunnerparity path on the shared SGT-009 / T-015 scenario - Scenario:
data/scenarios/multi_01_end_to_end_fault_response.json - Model: self-hosted
openai/Llama-3.1-8B-Instructthrough local vLLM on Insomnia - Current reachable PR branch: smoke-fix branch for
#104 - Historical Slurm-recorded SHAs: these jobs ran before the Apr 26
author/committer attribution rewrite. The run
meta.jsonfiles therefore record pre-rewrite hashes that are no longer reachable from the remote branch: Cell A9541e2661111daa14eb4d99f46d30bdc03681114, Cell Ba10d092d374309f45d282c7f7aec71a7fa8d11df, upstream paritye43cba33c7d78cf17390ec65bd82aeb4a9ebbe10. The rewritten PR branch above preserves the smoke-fix code/doc lineage and is the checkout target for review/merge. - Cell A config:
configs/aat_direct_smoke.env - Cell A run id / Slurm job id:
8962310_aat_direct_smoke_104 - Cell A primary artifacts: live artifacts in the shared Insomnia checkout:
benchmarks/cell_A_direct/raw/8962310_aat_direct_smoke_104/meta.jsonbenchmarks/cell_A_direct/raw/8962310_aat_direct_smoke_104/harness.logbenchmarks/cell_A_direct/raw/8962310_aat_direct_smoke_104/latencies.jsonlbenchmarks/cell_A_direct/raw/8962310_aat_direct_smoke_104/2026-04-25_A_llama-3-1-8b-instruct_agent_as_tool_direct_multi_01_end_to_end_fault_response_run01.jsonbenchmarks/cell_A_direct/config.json,benchmarks/cell_A_direct/summary.json
- Cell B config:
configs/aat_mcp_baseline_smoke.env - Cell B run id / Slurm job id:
8969519_aat_mcp_baseline_smoke_104 - Cell B primary artifacts: live artifacts in the shared Insomnia checkout:
benchmarks/cell_B_mcp_baseline/raw/8969519_aat_mcp_baseline_smoke_104/meta.jsonbenchmarks/cell_B_mcp_baseline/raw/8969519_aat_mcp_baseline_smoke_104/harness.logbenchmarks/cell_B_mcp_baseline/raw/8969519_aat_mcp_baseline_smoke_104/vllm.logbenchmarks/cell_B_mcp_baseline/raw/8969519_aat_mcp_baseline_smoke_104/latencies.jsonlbenchmarks/cell_B_mcp_baseline/raw/8969519_aat_mcp_baseline_smoke_104/2026-04-26_B_llama-3-1-8b-instruct_agent_as_tool_baseline_multi_01_end_to_end_fault_response_run01.jsonbenchmarks/cell_B_mcp_baseline/config.json,benchmarks/cell_B_mcp_baseline/summary.json
- Upstream parity config:
configs/aat_mcp_baseline_upstream_smoke.env - Upstream parity run id / Slurm job id:
8970383_aat_mcp_baseline_upstream_smoke_104 - Upstream parity primary artifacts: live artifacts in the shared Insomnia checkout:
benchmarks/cell_B_mcp_baseline/raw/8970383_aat_mcp_baseline_upstream_smoke_104/meta.jsonbenchmarks/cell_B_mcp_baseline/raw/8970383_aat_mcp_baseline_upstream_smoke_104/harness.logbenchmarks/cell_B_mcp_baseline/raw/8970383_aat_mcp_baseline_upstream_smoke_104/vllm.logbenchmarks/cell_B_mcp_baseline/raw/8970383_aat_mcp_baseline_upstream_smoke_104/latencies.jsonlbenchmarks/cell_B_mcp_baseline/raw/8970383_aat_mcp_baseline_upstream_smoke_104/2026-04-26_B_llama-3-1-8b-instruct_agent_as_tool_baseline_multi_01_end_to_end_fault_response_run01.json
- Upstream parity repeat run id / Slurm job id:
8970468_aat_mcp_baseline_upstream_smoke_104 - Upstream parity repeat primary artifacts: live artifacts in the shared Insomnia checkout:
benchmarks/cell_B_mcp_baseline/raw/8970468_aat_mcp_baseline_upstream_smoke_104/meta.jsonbenchmarks/cell_B_mcp_baseline/raw/8970468_aat_mcp_baseline_upstream_smoke_104/harness.logbenchmarks/cell_B_mcp_baseline/raw/8970468_aat_mcp_baseline_upstream_smoke_104/vllm.logbenchmarks/cell_B_mcp_baseline/raw/8970468_aat_mcp_baseline_upstream_smoke_104/latencies.jsonlbenchmarks/cell_B_mcp_baseline/raw/8970468_aat_mcp_baseline_upstream_smoke_104/2026-04-26_B_llama-3-1-8b-instruct_agent_as_tool_baseline_multi_01_end_to_end_fault_response_run01.json
What this proves:
- Cell A and Cell B now run through the same OpenAI Agents SDK loop with only the tool source changed: direct callables for Cell A, MCP stdio servers for Cell B
- the AaT runner reaches local vLLM, uses the pinned AOB prompt, and emits the canonical benchmark artifact set for both smoke cells
- Cell A completed
1 / 1withrun_status: "success", wall-clock latency 12.09 s, and 4 tool calls - Cell B completed
1 / 1withrun_status: "success", wall-clock latency 91.78 s, and 4 MCP tool calls after all four Smart Grid MCP servers bootstrapped and initialized - the Cell B smoke specifically validates the local-vLLM compatibility fixes:
explicit LiteLLM base URL/API key wiring, vLLM auto tool choice with
llama3_json, warmed.venv-insomniaMCP server launch, 120 s MCP initialize timeout, andparallel_tool_calls=falsefor sequential tool-call turns - the upstream parity smoke drove AssetOpsBench's
OpenAIAgentRunnerPython API end-to-end against the same Smart Grid MCP servers and scenario: SlurmCOMPLETED 0:0in00:11:18, benchmarkrun_status: "success",1 / 1scenario complete, 36.18 s benchmark latency, 30.14 s upstream runner duration, and 4 MCP tool calls - the repeat upstream parity smoke also succeeded end-to-end:
Slurm
COMPLETED 0:0in00:09:05, benchmarkrun_status: "success",1 / 1scenario complete, 31.48 s benchmark latency, and 4 MCP tool calls - the upstream parity harness reached MCP bootstrap, MCP initialize,
model requests, and tool execution: all four servers listed tools,
local vLLM served five
/v1/chat/completionscalls, and the trajectory calledget_sensor_readings,get_sensor_correlation,forecast_rul, andcreate_work_order - the Cell A / Cell B fairness guard now checks both model-visible tool names and per-tool parameter requiredness. Tool descriptions may still differ slightly between Python docstrings and FastMCP-derived schemas; the enforced contract is that the callable names and required argument surface match.
Caveats / follow-ups:
- these are one-scenario smoke proofs, not the full
#25Experiment 1 capture slice (multi_*.json, 3 trials, A/B/C) - Cell A was proven on an earlier branch tip before the final Cell B runtime hardening; the code-path changes after that point were MCP/vLLM compatibility fixes and did not change the direct tool surface
- Cell C is now proven separately in the Apr 30 entry above; at the time of these smoke proofs it was still pending optimized MCP readiness.
- the upstream parity proof uses AOB's
OpenAIAgentRunnerPython API rather than theopenai-agentCLI because the CLI cannot pass Smart Gridserver_paths; the wrapper keeps AOB's agent loop and patches only the MCP server launch envelope andparallel_tool_calls=falsefor Insomnia/local-vLLM compatibility
- Scope: end-to-end AssetOpsBench plan-execute run against all four Smart Grid MCP servers via Watsonx
- Branch / state: canonical
mainat the time of the run - Scenario:
data/scenarios/multi_01_end_to_end_fault_response.json(SGT-009, transformer T-015) - Model:
watsonx/meta-llama/llama-3-3-70b-instruct - Run name:
local-20260413-003914_pe_mcp_baseline_watsonx_smoke - W&B: 9d4442ja
- Primary artifacts (all committed in-tree):
benchmarks/cell_Y_plan_execute/raw/local-20260413-003914_pe_mcp_baseline_watsonx_smoke/meta.jsonbenchmarks/cell_Y_plan_execute/raw/local-20260413-003914_pe_mcp_baseline_watsonx_smoke/harness.logbenchmarks/cell_Y_plan_execute/raw/local-20260413-003914_pe_mcp_baseline_watsonx_smoke/latencies.jsonlbenchmarks/cell_Y_plan_execute/raw/local-20260413-003914_pe_mcp_baseline_watsonx_smoke/2026-04-13_Y_llama-3-3-70b-instruct_plan_execute_baseline_multi_01_end_to_end_fault_response_run01.jsonbenchmarks/cell_Y_plan_execute/config.json,benchmarks/cell_Y_plan_execute/summary.json
What this proves:
- the AssetOpsBench
plan-executeCLI successfully drove all four Smart Grid MCP servers end-to-end through the 8-tool-call sequence for SGT-009 - the benchmark wrapper produced canonical
benchmarks/cell_Y_plan_execute/raw/<run-id>/artifacts on the first committed proof run - wall-clock 93.6 s with the full LLM latency included;
run_status: "success",pass: 1,fail: 0 - this was the earliest committed in-tree proof of the benchmark-facing path and seeded the canonical Cell Y artifact layout that later runs reused
Caveats / follow-ups:
- single-scenario smoke; not a full grid
- uses WatsonX Llama-3.3-70B rather than the eventual Insomnia self-hosted Llama-3.1-8B path (those are the
#58and#115lane) - superseded as the canonical Cell Y snapshot by the Apr 21 PE + Self-Ask and Verified PE runs, but the raw artifacts remain committed as the earliest proof
- Scope: self-hosted Llama-3.1-8B benchmark-path validation on Insomnia
- Branch / state: PR
#115branch (not yet canonicalmainat the time of validation) - Key runtime shape:
--served-model-name Llama-3.1-8B-Instruct,--max-model-len 32768, local vLLM OpenAI-compatible path - Primary artifacts: committed validation artifacts referenced from PR
#115
What this proves:
- the long-context benchmark-facing serve path worked on an Insomnia A6000 node
- the benchmark path needed the served-model-name / OpenAI-client alignment
- the successful proof used the longer
32768context lane rather than the lighter8192smoke-path default
Caveats / follow-ups:
- the validated shape still needed to be folded back into shared scripts/docs on canonical history
- startup-time expectations from that run informed the later timeout cleanup
- Scope: repo-local PE + Self-Ask runner on Insomnia
- Branch / git SHA: historical pre-accounting-fix branch state (around
0591c75, pre-rebase) - Config:
configs/example_pe_self_ask.env - Run id / Slurm job id:
8850716_pe_self_ask_mcp_baseline_smoke - W&B:
y42u88h3 - Primary artifacts: historical live-run artifacts in the Insomnia worktree + W&B
y42u88h3
What this proves:
- the repo-local Self-Ask PE runner executed end-to-end on Insomnia
- local
vllm==0.19.0, Smart Grid MCP servers, LiteLLM/OpenAI-compatible local serving, and WandB upload all worked together in one live run
Caveats / follow-ups:
- this was an integration proof, not yet a clean method-quality proof
- one scenario still ended with a terminal failed step (
Unknown server 'none') even though the benchmark wrapper counted the run as completed - that accounting bug is now fixed on the branch; rerun after the fix is required before treating this as final PR evidence
- Scope: repo-local Verified PE runner on Insomnia
- Branch / git SHA: historical pre-accounting-fix branch state (around
0591c75, pre-rebase) - Config:
configs/example_verified_pe.env - Run id / Slurm job id:
8851966_verified_pe_mcp_baseline_smoke - W&B:
0v3a5jqi - Primary artifacts: historical live-run artifacts in the Insomnia worktree + W&B
0v3a5jqi
What this proves:
- the repo-local Verified PE workflow also executes end-to-end on Insomnia with live verifier / retry behavior
- the runtime stack is the same working local-serving path as the PE + Self-Ask run
Caveats / follow-ups:
- this run also happened before the benchmark-wrapper success-accounting fix
- the raw scenario outputs show semantic failures even though the wrapper summary reported
pass=2, so rerun on the fixed branch is required
- Scope: repo-local PE + Self-Ask runner on Insomnia
- Branch / git SHA: verified-PE/Self-Ask proof branch at
3a03ab83b7714c1d0f3aed2bc4899ef63fe5511c - Config:
configs/example_pe_self_ask.env - Run id / Slurm job id:
8857842_pe_self_ask_mcp_baseline_smoke - W&B: otkt77pj
- Primary artifacts:
- committed snapshot:
benchmarks/cell_Y_plan_execute/config.json - committed snapshot:
benchmarks/cell_Y_plan_execute/summary.json - live raw artifacts: archived in the Insomnia worktree under run id
8857842_pe_self_ask_mcp_baseline_smoke
- committed snapshot:
What this proves:
- the repo-local PE + Self-Ask runner reached a full
2 / 2smoke success on the two multi-domain scenarios on the rebased post-#115branch - the live path was clean end-to-end: local vLLM, LiteLLM/OpenAI-compatible serving, Smart Grid MCP servers, benchmark wrapper, and WandB upload
- the committed
config.json/summary.jsonsnapshot now gives the PR an in-tree proof surface without requiring the full raw log bundle in git
Caveats / follow-ups:
- the full raw logs and per-scenario JSONs are intentionally not committed in this branch; they remain archived on Insomnia and externally reflected in W&B
- earlier
8854783_pe_self_ask_mcp_baseline_smokeremains useful historical evidence, but8857842is the committed snapshot aligned to the current rebased branch state
- Scope: repo-local Verified PE runner on Insomnia
- Branch / git SHA: verified-PE/Self-Ask proof branch at
3a03ab83b7714c1d0f3aed2bc4899ef63fe5511c - Config:
configs/example_verified_pe.env - Run id / Slurm job id:
8857843_verified_pe_mcp_baseline_smoke - W&B: x65ej9e0
- Primary artifacts:
- committed snapshot:
benchmarks/cell_Z_hybrid/config.json - committed snapshot:
benchmarks/cell_Z_hybrid/summary.json - live raw artifacts: archived in the Insomnia worktree under run id
8857843_verified_pe_mcp_baseline_smoke
- committed snapshot:
What this proves:
- the repo-local Verified PE runner reached a full
2 / 2smoke success on the rebased post-#115branch - verifier-time prompt overflows, summarization overflows, and oversized execution-context recycling are all fixed enough for a clean live proof
- the committed
config.json/summary.jsonsnapshot now gives the PR an in-tree proof surface for the Verified PE lane as well
Caveats / follow-ups:
- this is the current authoritative Verified PE smoke snapshot for the PR
- the full raw logs and per-scenario JSONs are intentionally not committed in this branch; they remain archived on Insomnia and externally reflected in W&B
- Scope: full Experiment 1 Cell A (Agent-as-Tool, direct in-process tools) and Cell B (Agent-as-Tool, MCP baseline) capture across the canonical multi-domain scenario set, executed sequentially in one Slurm allocation via
scripts/run_exp1_ab_capture.sh. Includes the Apr 21 instrumentation validation: WandB run linkage, nvidia-smi GPU timeline, and a vLLM torch-profiler trace per cell. - Branch / git SHA:
aaron/exp1-ab-capturerooted ated136e8(committed via PR#130) - Slurm job id:
8979314 - Slurm state / elapsed:
COMPLETED 0:0in00:10:31 - Capture script:
scripts/run_exp1_ab_capture.sh - Configs:
- Cell A:
configs/aat_direct.env - Cell B:
configs/aat_mcp_baseline.env - both with
TORCH_PROFILE=1,LAUNCH_VLLM=1,ENABLE_WANDB=1,ENABLE_SMARTGRID_SERVERS=1(Cell B only)
- Cell A:
- Scenario set:
data/scenarios/multi_*.json(canonical multi-domain pack; 2 scenarios × 3 trials per cell) - Model: self-hosted
openai/Llama-3.1-8B-Instructvia local vLLM on Insomnia (vllm==0.19.0, FP16, max_model_len=8192)
benchmarks/cell_A_direct/raw/8979314_aat_direct/—meta.json(withwandb_run_urland repo-root-relativeprofiling_dir),summary.json(6 / 6scenarios, mean12.19s,run_status: success,git_sha: ed136e8),latencies.jsonl,harness.log,vllm.log, per-trial JSONs, replay JSONs.benchmarks/cell_A_direct/config.json+summary.json(cell-level rolled-up).- Profiling (gitignored):
profiling/traces/8979314_cell_a/nvidia_smi.csv— GPU util / memory / power timeline at 1 Hz, 355 sample rowsprofiling/traces/8979314_cell_a/nvidia_smi.stderr.log— sampler error sidecar (empty for this run)profiling/traces/8979314_cell_a/capture_meta.jsonprofiling/traces/8979314_aat_direct_torch/*.pt.trace.json.gz— PyTorch profiler Chrome trace from vLLM/start_profile+/stop_profile
- WandB run: https://wandb.ai/assetopsbench-smartgrid/assetopsbench-smartgrid/runs/vq976ljq
- WandB Artifact:
profiling-vq976ljq(gpu-util mean16.9 %/ max100 %, memory used mean23.8 GiB/ max42.1 GiB, power mean95.7 W/ max273.3 W)
benchmarks/cell_B_mcp_baseline/raw/8979314_aat_mcp_baseline/— same shape;summary.jsonis6 / 6, mean13.38s,run_status: success,git_sha: ed136e8.benchmarks/cell_B_mcp_baseline/config.json+summary.json.- Profiling (gitignored):
profiling/traces/8979314_cell_b/nvidia_smi.csv(249 sample rows),nvidia_smi.stderr.log(empty),capture_meta.json,profiling/traces/8979314_aat_mcp_baseline_torch/*.pt.trace.json.gz. - WandB run: https://wandb.ai/assetopsbench-smartgrid/assetopsbench-smartgrid/runs/qejvnoug
- WandB Artifact:
profiling-qejvnoug
- Cell A and Cell B run through the same OpenAI Agents SDK loop with only the tool source changed. (Cell B − Cell A) wall-clock =
13.38 − 12.19 = 1.20 smean is what Notebook 02 will treat as MCP transport overhead. - All three Apr 21 instrumentation streams produced artifacts and link back to benchmark run metadata:
- WandB:
wandb_run_urlis in bothmeta.jsons;log_profiling_to_wandb.pyattachednvidia_smi.csv+capture_meta.jsonas a WandB Artifact and pushed gpu-util / memory / power summary intowandb.run.summary.meta.json:profiling_diris repo-root-relative (profiling/traces/8979314_cell_{a,b}), portable across team / personal-scratch checkouts. - nvidia-smi: non-empty CSV timelines for both cells with 355 / 249 sample rows respectively under
profiling/traces/8979314_cell_{a,b}/. Samplerstderr.logsidecar is empty for both, confirming no transient nvidia-smi failures. - PyTorch / vLLM torch profiler: non-empty
*.pt.trace.json.gzunderprofiling/traces/8979314_aat_{direct,mcp_baseline}_torch/, captured via the vLLM 0.19.0--profiler-configCLI flag (profiler=torch,torch_profiler_dir=...) +scripts/replay_scenarios.shwhile vLLM was still alive in each cell's job phase.
- WandB:
Three fixes to the team's instrumentation infrastructure that the first run attempts on Apr 26 surfaced (9430b09), plus two more added during PR review iteration (a7de839, ed136e8).
scripts/aat_runner.py:266—json.dumps(..., default=str)in_write_output. Tool results sometimes carrypandas.Timestamp/numpy.datetime64objects that vanillajson.dumpsrejects;default=strmakes the encoder fall back tostr(). Crashed 1 trial in8978161Cell A; fix delivers 6/6.profiling/scripts/capture_around.sh:133— post-run WandB uploader picks callerPYTHON_BIN→.venv-insomnia/bin/python→python3instead of barepython3. Insomnia's system Python 3.9 has nowandb; barepython3silently dropped the Artifact upload. With the fix,capture_aroundfinishesrc=0and the artifact attaches to the WandB run.scripts/run_experiment.sh:753-758— vLLM 0.19.0 droppedVLLM_TORCH_PROFILER_DIR(logsUnknown vLLM environment variable detected). Profiling now requires the--profiler-configCLI flag (profiler=torch, absolutetorch_profiler_dir). The patched block builds the JSON config and appends toVLLM_SERVER_ARGSso/start_profileis registered in the FastAPI app.profiling/scripts/log_profiling_to_wandb.py— relativizeprof_diragainst the repo root before writing intometa.jsonand WandB config. Fixes Alex's High #3: prior runs leaked the personal-scratch absolute path/insomnia.../af3623/exp1-clone/...into committed artifacts.profiling/scripts/sample_nvidia_smi.sh— disablepipefailfor the header-write line. The pipenvidia-smi --format=csv | head -1 > FILEwas sending SIGPIPE to nvidia-smi (exit 141) which underset -euo pipefailkilled the script before the sample loop started — silently producing 1-line (header-only) CSVs in every prior capture (8962310,8969519,8978297,8979215). Loop also getsset +eplus a<name>.stderr.logsidecar to survive transient nvidia-smi failures during sampling. This is the actual root cause behind Alex's Critical #1.
- Cell C (MCP optimized) was not part of this Apr 26 A/B capture; the Apr 30 Cell C entry above supersedes the old optimized-lane gate.
profiling/traces/is gitignored. Paths above point at the live Insomnia checkout. The WandB Artifact uploads are the portable copies.- Cell A's torch-profiler replay pass got
pass=1 fail=1:multi_01_end_to_end_fault_responsehitContextWindowExceededError(8193 vs 8192 tokens). Non-fatal — the main scenario loop's 6/6 is the canonical capture; the replay only feeds the second torch-profiler trace. BumpingMAX_MODEL_LENto 16384 in the configs is a follow-up tweak, not in this PR's scope. - Run was executed from a personal scratch clone at
/insomnia001/depts/edu/users/af3623/exp1-clone/because the team-shared checkout's.git/objectshad perm issues for non-wax1writers. Personal clone symlinksmodels/and.venv-insomnia/from the shared checkout for storage efficiency. - The shared
.venv-insomniawas extended viauv pip install -r requirements-insomnia.txtto addopenai-agents==0.14.5+griffelib+types-requests, and refreshedwebsockets16.0 → 15.0.1. Pinged Tanisha. - Earlier
8978297and8979215attempts were uncommitted / removed in this PR's history because they suffered from the bugs the rerun fixes.