2026 · lightning
Evaluating AI Agent Tooling
The evaluation harnesses are currently being created to test Agentic tooling. There’s quite a bit of traction around evaluating custom built agents and agentic workflows. However, we’re now seeing tooling orchestrations built on top of Agents. Think “I want to write a bunch of prompts and instruction files and put them into my repo. How then do you judge if you’re making things better or not.” Also, are your instructions good for Github Copilot and bad for Claude? If you’re an AI architect you need to have answers to these questions. There’s a new tool out called Harbor built by laude institute, which is showing promise. It’s built on top of terminal-bench. I’m currently working on a project to automate running VSCode, starting off with SWE-Bench Verified metrics. I’d like to go over SWEBench, what it does and how it’s helpful, writing your own evaluations, and what it’s like doing this research