Finish the truncation change: the log, the final frame, the README and the changelog - #104
Merged
Merged
Conversation
Two places that still read terminated from when that one flag also covered the clock. Splitting the two causes was right; neither of these moved with it. The log said "Terminated: Reached max time" for the case the flags now deliberately call truncation. That line is the one thing an operator reads to find out why a run stopped, so it should not name the wrong ending. The renderer drew a frame every 0.1 s of simulation time or when the episode ended, and the ending it checked was terminated alone. While that covered the clock, an episode running out of horizon always drew its last frame. Afterwards it only did so when the final step happened to land on the cadence. Scenario 0 truncates at step 9999 against a cadence of 10, so it did not, and scenario 1 ends this way normally. The existing console test pinned the old wording, so it moves too. Verified by mutation: both the old wording and the old render gate fail. The frame test also asserts that the final step misses the cadence, since otherwise the old gate would have drawn it anyway and the assertion would hold either way. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
Merged
Sweeping for everything that still read terminated from when that one flag also covered the clock turned up two more than the code. The README described step() as returning "the new observations, reward, termination flag, and additional info", which is four values and one flag, and said the episode ends when the time limit is reached or the rocket lands, as though those were the same outcome. It is the document a competitor writes their loop from, so it was the worst place for this to be stale. It now names both flags, says which cause each one is, and shows the loop. The changelog had no entry at all. Added under Changed as breaking for every agent loop, alongside the observation copies, and the four behaviour fixes from this wave under Fixed. The README snippet is not a .py file, so the invariant test could not see it. It reads the fenced python blocks out of the README and puts any loop it finds through the same guard and unpack checks as the runners, plus a check that it found one at all, since a snippet that stops being tagged python would otherwise leave the check holding over nothing. Verified by mutation: the README showing the old loop, the README discarding truncated, and the snippet ceasing to be python all fail. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com>
2 tasks
zuorenchen
added a commit
that referenced
this pull request
Aug 27, 2026
* Refuse to pack a submission from a run that never finished build_submission_payload reads the score and the trajectories straight off the environment, and nothing said whether the episode had reached an ending. A run stopped part way through produced a submission that looked exactly like a complete one, carrying whatever score it had reached by then. Measured on scenario 0: five steps in, with the rocket still on the pad, the payload came out fully formed and leaderboard_info had no field that mentioned the run being unfinished. The environment now records how the episode ended, reset clears it, the payload carries it along with the step count, and pack_for_submission refuses a run that never got there. Which of terminated and truncated should score is #104 and is not decided here; this only asks whether the episode reached an ending at all. The fake environments in the submission tests gain the same two fields. They document themselves as the attributes pack_for_submission reads, and that set grew. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Clean up comments --------- Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> Co-authored-by: zuorenchen <zuorenchen@m110.nthu.edu.tw> Co-authored-by: ZuoRen Chen <180084773+zuorenchen@users.noreply.github.com>
zuorenchen
added a commit
that referenced
this pull request
Aug 29, 2026
* Write down that the launch step's control fields are not applied Closes #80, as decided there: document the behaviour and leave the implementation alone. The rocket is still on the rail on the step that launches, so `tvc`, `throttle` and `roll` would not change where it goes. That step builds the flight from the launch attitude and the step after it is the first that applies them. Two lines, one in the README where the action space is described and one in `step()` where a reader would ask. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Say why an action field was ignored The catch around the conversion is broad on purpose, and the log line it fed said only that a field was ignored. A competitor sending a string and a defect in the four lines above it produced the same sentence, so the defect was invisible. The reason travels with the field name now: Step 42: ignoring tvc, which the environment cannot use: ValueError: could not convert string to float: 'x' (1 such steps for that field) check_action still answers with a sorted list of names, since that is what an agent calls. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Record what landed after v0.1.1 The Unreleased section was empty again within the hour, which is the drift #134 existed to stop. Two entries, for #131 and #126. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Put a placeholder where the example carried a real secret example_eval_cfg.yaml and the README both shipped a working team_secret. A credential does not belong in an example, and a competitor has to replace it with their own anyway, so the file now says what to paste instead of shipping something that already works. No behaviour changes. The secret is only ever copied into the packed submission, and nothing reads its value. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Fix inclination indication in readme fig (#153) * ENH: Add wall time limit (#154) * Add wall time limit * Update tests * Update the suggestions from PR review * Start the episode clock on the same one step reads, and pin it (#156) * Start the episode clock on the same one step reads reset() started it on time.time() while step() measured with time.monotonic(), so the difference was about -1.8e9 and no episode could reach any limit. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Pin the wall clock limit with the test that would have caught it Nothing in tests/ mentioned max_wall_time, and a limit that can never fire looks exactly like a limit nobody reached. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> --------- Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> --------- Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> Co-authored-by: 秀吉 <84045975+thc1006@users.noreply.github.com> * Refuse to pack a submission from a run that never finished (#142) * Refuse to pack a submission from a run that never finished build_submission_payload reads the score and the trajectories straight off the environment, and nothing said whether the episode had reached an ending. A run stopped part way through produced a submission that looked exactly like a complete one, carrying whatever score it had reached by then. Measured on scenario 0: five steps in, with the rocket still on the pad, the payload came out fully formed and leaderboard_info had no field that mentioned the run being unfinished. The environment now records how the episode ended, reset clears it, the payload carries it along with the step count, and pack_for_submission refuses a run that never got there. Which of terminated and truncated should score is #104 and is not decided here; this only asks whether the episode reached an ending at all. The fake environments in the submission tests gain the same two fields. They document themselves as the attributes pack_for_submission reads, and that set grew. Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> * Clean up comments --------- Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> Co-authored-by: zuorenchen <zuorenchen@m110.nthu.edu.tw> Co-authored-by: ZuoRen Chen <180084773+zuorenchen@users.noreply.github.com> * Fix pylints * docs: make README setup commands cross-platform (#138) * docs: make README setup commands cross-platform * Potential fix for pull request finding --------- Co-authored-by: 秀吉 <84045975+thc1006@users.noreply.github.com> Co-authored-by: ZuoRen Chen <180084773+zuorenchen@users.noreply.github.com> * Revert "Put a placeholder where the example carried a real secret" (#159) * Fix: ActiveRocketPy bug fix (#157) * Update ActiveRocketPy submodule to fix TVC and acclerometer bug * Fix the incorrect accelerometer model following the updates in ActiveRocketPy * Regenerate scenario 0 and 1 baselines * Bump submission version to 2 (#160) * Update changelog for v0.2.0 --------- Signed-off-by: thc1006 <84045975+thc1006@users.noreply.github.com> Co-authored-by: thc1006 <84045975+thc1006@users.noreply.github.com> Co-authored-by: William Mou <william.mou1024@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Finishes the split between
terminatedandtruncated. That split was right; several places still readterminatedfrom when that one flag also covered the clock, and none of them moved with it.I swept for all of them rather than fixing the two that were reported.
The environment
Running out of horizon is truncation now, and this line is the one thing an operator reads to find out why a run stopped.
Frames are drawn every 0.1 s of simulation time, or when the episode ends. While the clock counted as termination, an episode running out of horizon always drew its final frame. Afterwards the gate only fires on
terminated, so a truncated episode draws one only if its last step happens to land on the cadence. Scenario 0 truncates at step 9999 against a cadence of 10, so it does not. Scenario 1 ends this way normally.The README, which is where a competitor writes their loop from
It described
env.step()as returning "the new observations, reward, termination flag, and additional info", which is four values and one flag, and said the episode ends when the time limit is reached or the rocket hits the ground, as though those were one outcome. It now names both flags, says which cause each is, and shows the loop.The changelog
Had no entry for any of this. Added under Changed as breaking for every agent loop, alongside the observation copies from #99, and the four behaviour fixes from this wave under Fixed.
What I checked, and what was already fine
evaluate.pydoc/examples/run_env_agent.pydoc/examples/test_navigation_agent.pyevaluate_scenario_colab.ipynb"Terminated: Rocket flight finished"Tests
The two environment fixes are covered in
tests/test_episode_lifecycle.py, driving a no-launch episode to the horizon. The frame test also asserts that the final step misses the render cadence, since otherwise the old gate would have drawn it anyway and the assertion would have held against either version.The README snippet is not a
.pyfile, so the invariant test added in #102 could not see it. It now reads the fencedpythonblocks out of the README and puts any loop it finds through the same guard and unpack checks as the runners, plus a check that it found one at all, since a snippet that stopped being taggedpythonwould otherwise leave that holding over nothing.The existing console-logging test pinned the old wording, so it moves too.
Verified by mutation, all five caught:
while not terminated:truncatedpythonLocal CI green on a clean tree: ruff check, ruff format,
uv lock --check, 315 passed withBPC_RUN_SLOW_TESTS=1.