Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments
Abstract
A scalable pipeline called TerminalTraj addresses challenges in creating high-quality terminal trajectories for training agentic models by filtering repositories, generating Docker-aligned task instances, and synthesizing executable agent trajectories across multiple domains.
Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse domains. However, constructing such data at scale remains challenging due to two key requirements: \emph{Executability}, since each instance requires a suitable and often distinct Docker environment; and \emph{Verifiability}, because heterogeneous task outputs preclude unified, standardized verification. To address these challenges, we propose TerminalTraj, a scalable pipeline that (i) filters high-quality repositories to construct Dockerized execution environments, (ii) generates Docker-aligned task instances, and (iii) synthesizes agent trajectories with executable validation code. Using TerminalTraj, we curate 32K Docker images and generate 50,733 verified terminal trajectories across eight domains. Models trained on this data with the Qwen2.5-Coder backbone achieve consistent performance improvements on TerminalBench (TB), with gains of up to 20\% on TB~1.0 and 10\% on TB~2.0 over their respective backbones. Notably, TerminalTraj-32B achieves strong performance among models with fewer than 100B parameters, reaching 35.30\% on TB~1.0 and 22.00\% on TB~2.0, and demonstrates improved test-time scaling behavior. All code and data are available at https://github.com/Wusiwei0410/TerminalTraj.
Community
This is a repo for paper "Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments"
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- DockSmith: Scaling Reliable Coding Environments via an Agentic Docker Builder (2026)
- TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents (2026)
- MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering (2026)
- SWE-World: Building Software Engineering Agents in Docker-Free Environments (2026)
- SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training (2026)
- Endless Terminals: Scaling RL Environments for Terminal Agents (2026)
- ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment:
@librarian-bot
recommend
Models citing this paper 3
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper