Skip to main content
AdOpenFree logoPromote your productReach more potential users and drive product growth and revenue.Advertise
Favicon of Harbor

Harbor

Free Listing

Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing agents and language models in container environments. It supports arbitrary agents, custom benchmarks, parallel cloud experiments, and rollouts for RL

Visit harbor-framework/harbor

Overview

Harbor is a framework from the creators of Terminal-Bench for specifying sandboxed agent tasks for evaluation and optimization. It can evaluate arbitrary agents such as Claude Code, OpenHands, and Codex CLI, build and share benchmarks and environments, run experiments in parallel across thousands of environments, and generate rollouts for RL optimization.

Key Features

  • Evaluate arbitrary agents like Claude Code, OpenHands, Codex CLI, and more.
  • Build and share your own benchmarks and environments.
  • Conduct experiments in thousands of environments in parallel through providers like Daytona, Modal, LangSmith, Blaxel, Novita Sandbox, Tensorlake, and Runta.
  • Generate rollouts for RL optimization.
  • Serves as the official harness for Terminal-Bench-2.0.

Use Cases

  • Running benchmarks such as Terminal-Bench-2.0 locally with Docker or on cloud providers.
  • Evaluating arbitrary agents and language models on supported datasets.
  • Exploring third-party benchmarks like SWE-Bench and Aider Polyglot.
  • Generating rollouts for RL optimization.

Getting Started

  • Install Harbor with uv tool install harbor or pip install harbor.
  • Run a benchmark such as: harbor run --dataset terminal-bench@2.0 --agent claude-code --model anthropic/claude-opus-4-1 --n-concurrent 4.
  • Set ANTHROPIC_API_KEY before running the Anthropic example.
  • Use harbor run --help to see supported agents and options.
  • Use harbor datasets list to explore supported third-party benchmarks.

Deployment & Requirements

  • Local benchmark runs use Docker.
  • Cloud runs can use providers such as Daytona via the --env flag.
  • Set ANTHROPIC_API_KEY for Anthropic models and DAYTONA_API_KEY for Daytona cloud runs.

Before You Adopt

  • License: Apache-2.0. Review its terms before using, modifying, or distributing the project.
  • Local benchmark runs require Docker.
  • Cloud-parallel experiments depend on third-party providers such as Daytona, Modal, LangSmith, Blaxel, Novita Sandbox, Tensorlake, and Runta.
  • The Anthropic example requires an ANTHROPIC_API_KEY environment variable.
  • Cloud runs on Daytona require a DAYTONA_API_KEY environment variable.

Comments

Sign in to leave a comment.

More like Harbor

Favicon of OpenHuman

OpenHuman

Free ListingStars: 40.1K

Open-source personal AI that runs on your laptop

Agent Frameworks

OpenHuman is a free, open-source personal AI for Mac, Windows and Linux that keeps data local and orchestrates fleets of agents. Built in Rust, it is positioned as a fast, efficient open-source agent harness.

Favicon of Browser Harness

Browser Harness

Free ListingStars: 18.1K

Self-healing browser agents built directly on CDP

Agent Frameworks

Browser Harness is a thin, self-healing harness for browser agents built directly on CDP. Agents can edit their own helpers mid-task, so the harness improves with every task and enables LLMs to complete browser tasks.

Favicon of DeepSeek Harness

DeepSeek Harness

Free ListingStars: 237.5K

Everything is a plugin.

Agent Frameworks

DeepSeek Harness (dsh) is an open-source agent harness in developer preview. Every agent capability is implemented as a composable Cordis plugin for models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI.