Reach more potential users and drive product growth and revenue.
AdvertiseTools tagged with "Evaluation"
Warp
Free ListingStars: 65.2KOpen infrastructure for cloud software factories
Warp is an agentic development environment, born out of the terminal, offering open infrastructure to build, measure, and interact with agents across the SDLC so teams ship more and spend less.
Langfuse
Free ListingStars: 35.1KOpen Source Agent Evals & Observability
Open source platform for tracing, evaluating, and improving AI agents. Use production data to understand behavior, collaborate on fixes, and ship better quality at lower cost and latency.
Agent Development Kit (ADK)
Free ListingStars: 21.6KOpen-source, code-first Python framework for AI agents
An open-source, code-first Python framework for building, evaluating, and deploying sophisticated AI agents with flexibility and control. ADK is model-agnostic, deployment-agnostic, and compatible with other frameworks.
Harbor
Free ListingStars: 5.6KEvaluate and optimize sandboxed agents and models.
Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing agents and language models in container environments. It supports arbitrary agents, custom benchmarks, parallel cloud experiments, and rollouts for RL

