A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and auditable.
-
Updated
Sep 13, 2026 - Python
A research framework for principled agent self-improvement under frozen evaluators and declared mutation boundaries, recording verifiable lineage to make it reproducible and auditable.
A plugin for your agentic framework that optimizes code using the GEPA algorithm (Genetic-Pareto LLM-driven search).
Synth Python SDK for Managed Research, Research Factory, and GEPA/GELO optimizer workflows.
Claude Code for DSPy: Comprehensive CLI to Optimize Your DSPy Code. our AI-Powered DSPy Development Assistant
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
CLI text optimizer built on GEPA. Uses Agentic Coding CLI's as mutator and observer -- no api keys required
GEPAzilla: open-source GEPA prompt optimizer with datasets, scorers, and telemetry.
Local code search for AI agents: six fast, purpose-built tools that return ranked answers, not raw grep. Because maybe grep isn't all you need... 🍬
Evolving agent harnesses: a research program on how far N orchestrated calls of a small model can rival a frontier model. We evolve the harness (structure + prompts) with reflective optimizers + a verified-acceptance gate.
GEPA and GELO optimizer runbooks, SDKs, and hosted optimizer surfaces for Synth.
Self-evolve Gemini CLI instructions, commands, and skills via the gemini CLI itself — GA + GEPA/DSPy, with hard gates before apply.
Which host should serve your open model? Run your workload across pinned, verified providers and get one table: success, cost, latency, cache hits.
Local, private prompt optimization with DSPy and Ollama. Rewrites prompts, measures them against your own data, and evolves them from feedback with GEPA.
Production-ready boilerplate for building and automatically optimizing LangChain RAG applications. Implements three-layer architecture: Build (LangChain) → Measure (MLflow) → Optimize (GEPA + MEGA).
Prompt optimisation with GEPA: mine a compliance rubric from labelled decisions. 30% more violations caught, starting from a one-line prompt.
Open-source LLM eval workbench: generate verifiable prompts from parametric templates, compare any model on text or image, settle quality with blind preference, evolve with GEPA, and serve self-hosted vLLM on your GPU.
Self-improvement feedback loop for AI agent skills — analyzes past sessions, drafts structured improvement proposals, and gates auto-apply behind an evaluation framework. Host-agnostic (Hermes, Claude Code, and any HostAdapter).
A benchmark, alignment pipeline, and LLM-as-a-Judge for evaluating the clinical impact of ASR errors.
DSPy/GEPA-powered self-improvement plugin for Hermes Agent skills, memory, and evaluator prompts.
To associate your repository with the gepa topic, visit your repo's landing page and select "manage topics."