Skip to main content
AdOpenFree logoPromote your productReach more potential users and drive product growth and revenue.Advertise

Tools tagged with "Evaluation"

Favicon of Warp

Warp

Free ListingStars: 65.2K

Open infrastructure for cloud software factories

Coding Agents

Warp is an agentic development environment, born out of the terminal, offering open infrastructure to build, measure, and interact with agents across the SDLC so teams ship more and spend less.

Favicon of Langfuse

Langfuse

Free ListingStars: 35.1K

Open Source Agent Evals & Observability

Evaluation & Benchmarks

Open source platform for tracing, evaluating, and improving AI agents. Use production data to understand behavior, collaborate on fixes, and ship better quality at lower cost and latency.

Favicon of Agent Development Kit (ADK)

Agent Development Kit (ADK)

Free ListingStars: 21.6K

Open-source, code-first Python framework for AI agents

Agent Frameworks

An open-source, code-first Python framework for building, evaluating, and deploying sophisticated AI agents with flexibility and control. ADK is model-agnostic, deployment-agnostic, and compatible with other frameworks.

Favicon of Harbor

Harbor

Free ListingStars: 5.6K

Evaluate and optimize sandboxed agents and models.

Agent Frameworks

Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing agents and language models in container environments. It supports arbitrary agents, custom benchmarks, parallel cloud experiments, and rollouts for RL

Ad
OpenFree logoPromote your product

Reach more potential users and drive product growth and revenue.

Advertise