Harbor
Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing agents and language models in container environments. It supports arbitrary agents, custom benchmarks, parallel cloud experiments, and rollouts for RL
Visit harbor-framework/harborOverview
Harbor is a framework from the creators of Terminal-Bench for specifying sandboxed agent tasks for evaluation and optimization. It can evaluate arbitrary agents such as Claude Code, OpenHands, and Codex CLI, build and share benchmarks and environments, run experiments in parallel across thousands of environments, and generate rollouts for RL optimization.
Key Features
- Evaluate arbitrary agents like Claude Code, OpenHands, Codex CLI, and more.
- Build and share your own benchmarks and environments.
- Conduct experiments in thousands of environments in parallel through providers like Daytona, Modal, LangSmith, Blaxel, Novita Sandbox, Tensorlake, and Runta.
- Generate rollouts for RL optimization.
- Serves as the official harness for Terminal-Bench-2.0.
Use Cases
- Running benchmarks such as Terminal-Bench-2.0 locally with Docker or on cloud providers.
- Evaluating arbitrary agents and language models on supported datasets.
- Exploring third-party benchmarks like SWE-Bench and Aider Polyglot.
- Generating rollouts for RL optimization.
Getting Started
- Install Harbor with uv tool install harbor or pip install harbor.
- Run a benchmark such as: harbor run --dataset terminal-bench@2.0 --agent claude-code --model anthropic/claude-opus-4-1 --n-concurrent 4.
- Set ANTHROPIC_API_KEY before running the Anthropic example.
- Use harbor run --help to see supported agents and options.
- Use harbor datasets list to explore supported third-party benchmarks.
Deployment & Requirements
- Local benchmark runs use Docker.
- Cloud runs can use providers such as Daytona via the --env flag.
- Set ANTHROPIC_API_KEY for Anthropic models and DAYTONA_API_KEY for Daytona cloud runs.
Before You Adopt
- License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
- Local benchmark runs require Docker.
- Cloud-parallel experiments depend on third-party providers such as Daytona, Modal, LangSmith, Blaxel, Novita Sandbox, Tensorlake, and Runta.
- The Anthropic example requires an ANTHROPIC_API_KEY environment variable.
- Cloud runs on Daytona require a DAYTONA_API_KEY environment variable.