Skip to content

Add YYLO Benchmark to LLM Evaluation Tools - #217

Open
InsightFactoryAPP wants to merge 1 commit into
promptslab:mainfrom
InsightFactoryAPP:add-yylo-benchmark
Open

InsightFactoryAPP wants to merge 1 commit into
promptslab:mainfrom
InsightFactoryAPP:add-yylo-benchmark

Conversation

@InsightFactoryAPP

Copy link
Copy Markdown

YYLO Benchmark is a CLI for longitudinal evaluation of agent runs — task prompts execute in private fresh-repository attempt workspaces under deterministic commands and configurable LLM judges, with hash-linked receipts, manifests, and provenance retained as immutable evidence.

  • Two deliberately separate evaluation lanes: isolated v2 (default — private fresh-repository attempt workspaces) and an explicitly authorized governed lane with blinded per-step judging
  • Evaluator profiles support deterministic commands and configurable LLM judges; a required deterministic correctness or safety failure cannot be overridden by judge prose
  • Installable via npm as @yylo/benchmark (187 downloads last month); the companion @yylo/cli coding-agent orchestrator is at 794/month
  • Open source (MIT), pushed near-daily since January 2026

Sits next to DeepEval / InspectAI / EvalView as an evaluation-layer tool for agents — not a prompt library.

Disclosure: I am on the YYLO team; this PR was prepared with AI assistance. Happy to revise, trim, or re-file if the placement or format needs adjusting.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant