Task description
# tree_sitter_markdown_inline__05
Create an editable project in `/app/workspace` that reproduces the observed
behavior of the reference demo for `tree_sitter_grammar`.
The public observations in `/app/observations` were produced by running the
real software on the acquired parent demo. Use them as behavioral evidence, not
as files to copy into your answer.
This variant changes at least 5 parameter axes: inline, heading, link, code_fence, table.
The private judge uses hidden scenarios centered on `table`.
Build the underlying mechanism so that unseen parameter settings continue to
work, instead of branching only on the public case.
## Required Reconstruction Interface
In addition to building an editable project, provide `/app/workspace/reconstruct.py`.
The verifier will execute it as:
```bash
python3 /app/workspace/reconstruct.py < scenario.json
```
It must read one scenario JSON object from stdin, use `scenario["parameters"]`,
and print JSON to stdout. Follow the same output shape as the public
`oracle_output.json` files inside `/app/observations/reference_outputs.tar.gz`;
usually this means printing `{"summary": ...}`.
Parameter axes: `inline, heading, link, code_fence, table`
Public scenarios:
- `public_1` parameters: `{"code_fence": 5.1, "heading": 6.1, "inline": 6.2, "link": 1.3, "table": 6.9}`
- `public_2` parameters: `{"code_fence": 6.0, "heading": 1.1, "inline": 5.5, "link": 2.0, "table": 5.3}`
Use `/app/tools/feedback` while solving. It runs your `reconstruct.py` on the
public scenarios and reports which public oracle signatures match. Passing
public feedback is not sufficient by itself; hidden scenarios from the same
parameter space are used for final grading.
How this rewrite was made
- Expert. Base Qwen3.8-27B solved the task under continue-until-timeout. It never marked the task complete: it kept working until the 2-hour limit ended the run, and the work it left in place passed.
- Runbook. Base Qwen3.8-27B read the expert trajectory and wrote runbook_01: what to build, milestones with example commands, and checks. The leak judge checked it before it was used.
- Rewrite. Base Qwen3.8-27B ran the task as plain Terminus-2 with the runbook in a private channel. 4 of 4 replays passed, in 13, 15, 15, 34 steps; the one shown below took 34.
Trajectories
| Trajectory | Steps | Commands | Agent time | Flagged replies | Command mix |
|---|
exploreeditruninstallwaitother Category comes from the first line of each command.
Shortened runbook (runbook_01): a brief summary, not the actual runbook used
A short plan for rebuilding the tree-sitter Markdown demo: learn what the reference outputs look like, write a script that reproduces them, and check that it generalizes beyond the public examples.
This summary was shortened for publication; the actual runbook used for the rewrite is not shown.
Rows line up by step number, not by meaning. Hover a command to see all of it.