Skip to content

docs: foreground benchmark evidence - #13

Merged
mohammadrezwankhan merged 1 commit into
mainfrom
agent/accessible-benchmark-readme
Aug 11, 2026
Merged

docs: foreground benchmark evidence#13
mohammadrezwankhan merged 1 commit into
mainfrom
agent/accessible-benchmark-readme

Conversation

@mohammadrezwankhan

Copy link
Copy Markdown
Owner

Summary: Add a concise benchmark-results overview, explicit research-versus-deployment boundary, and the committed primary revision-2 figure with provenance links. Validation: 67 unit tests, five result-bundle audit checks, Markdown links, and diff check.

@mohammadrezwankhan
mohammadrezwankhan marked this pull request as ready for review August 11, 2026 19:45
Copilot AI lite review requested due to automatic review settings August 11, 2026 19:45
@mohammadrezwankhan
mohammadrezwankhan merged commit d71e697 into main Aug 11, 2026
4 checks passed
@mohammadrezwankhan
mohammadrezwankhan deleted the agent/accessible-benchmark-readme branch August 11, 2026 19:45

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds a concise “benchmark evidence” section to the top-level README to summarize revision-2 results and reinforce the research-only (non-deployment) scope, including a committed primary figure and provenance links for auditability.

Changes:

  • Added a “What this benchmark demonstrates” overview with key synthetic + historical outcome summaries.
  • Embedded the primary revision-2 results figure and linked it to the manifest and audit trail.
Suppressed comments (1)

README.md:39

  • The text alternates between “revision 2” and “revision-2”. For consistency (and easier searching), it would help to stick to the same spelling as the section heading (“revision 2”).
![Primary revision-2 result: synthetic paired differences and historical day-ahead block profit](results_revision2/figures/Figure_3_comparative_results.png)

*Primary result snapshot from the committed revision-2 bundle, generated by
the [`voltrl_benchmark.py`](voltrl_benchmark.py) command below. Panel A shows

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread README.md
Comment on lines +25 to +32
VoltRL is an auditable research benchmark under two explicit information
protocols, not a live-trading or deployment system. Across 30 regenerated
synthetic datasets, seasonal autoregressive (SARX) 24-hour MPC has the highest
mean annualized adjusted profit (EUR 5.13M; 95% bootstrap CI EUR 5.05–5.23M),
versus EUR 4.62M for the hour-aware finite-horizon MDP (EUR 4.51–4.73M).
Historical DK1/DK2 schedules are fixed before delivery and use realized prices
only for settlement; SARX reports EUR 0.129M and EUR 0.187M annualized adjusted
profit, respectively, after modeled degradation. These protocol-bound results
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants