Rubrics for evaluating AI agents

A repeatable, rubric-based way to evaluate how AI agents behave when they act on real systems through MCP tools — consistent scores instead of one-off manual checks.

  • AI agents
  • MCP
  • testing
  • LLM evaluation

narmaku.com

This site: a Jekyll landing page plus a Chirpy blog, built in one pipeline and served from Cloudflare's edge, mirrored on GitHub Pages.

  • Jekyll
  • Cloudflare Workers
  • GitHub Actions