Guides
Updated
Task-shaped walkthroughs for the work you actually do.
Each guide starts from a job and shows how to run it on Jetty: what goes in the runbook, what the runs show, and what to change when one misses. Start with the use cases to see the loop on another team's workload.
Build vs. Buy EvaluationWhat AI agent evaluation costs to build versus buy: the variables that matter, a framework for making the call, and how to test it against your own workload.Building AI EvaluationsA way to know, before you ship, whether your agent is actually doing the job it was built to handle at the level of quality you expect.Creating a BenchmarkPick one axis of comparison, write the runbook that defines the job, and run the task set against every candidate in parallel so the comparison is reproducible.Defining Agent ChecklistsKeep your agent on task with a checklist: the conditions that define done, ticked and rechecked inside every run, then investigated across runs and kept honest in production.Improving Agent RunbooksPull the latest runs with your coding agent, tighten the evals they expose, and let the optimize-runbook skill propose evidence-backed edits.ResearchPapers Jetty has contributed to, on agent benchmark standards, reproducible ML evaluations and AI accountability.Scheduling AI WorkloadsRun a runbook on a cron cadence so evals stay fresh as models and providers drift.Use CasesStories from teams running AI workloads on Jetty: what they set up, how parallel runs and evals changed the work, and what came out.