<p align=“center”> <picture> <source media=“(prefers-color-scheme: dark)” srcset=“assets/logo-dark.png”> Ponytail, the lazy senior dev </picture> </p>

<p align=“center”> <picture> <source media=“(prefers-color-scheme: dark)” srcset=“assets/logo-dark.png”> 马尾辫,那个懒惰的高级开发 </picture> </p>

<h1 align=“center”>Ponytail</h1>

<h1 align=“center”>马尾辫</h1>

<p align=“center”> He says nothing. He writes one line. It works. </p>

<p align=“center”> 他沉默不语。他写下一行代码。然后就能运行。 </p>

<p align=“center”> Stars Release npm Works with 20 agents MIT license </p>

<p align=“center”> 星标 版本 npm 兼容20个代理 MIT许可 </p>

<p align=“center”> <a href=“https://trendshift.io/repositories/50668” target=“_blank” rel=“noopener noreferrer”>DietrichGebert/ponytail | Trendshift</a> <a href=“https://trendshift.io/repositories/50668” target=“_blank” rel=“noopener noreferrer”>DietrichGebert/ponytail | Trendshift</a> </p>

<p align=“center”> <a href=“https://trendshift.io/repositories/50668” target=“_blank” rel=“noopener noreferrer”>DietrichGebert/ponytail | 趋势变化</a> <a href=“https://trendshift.io/repositories/50668” target=“_blank” rel=“noopener noreferrer”>DietrichGebert/ponytail | 趋势变化</a> </p>

<p align=“center”> ~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe

<sub>Measured on real Claude Code sessions editing a real open-source repo (FastAPI + React), against the same agent with no skill. ~54% is the mean across 12 feature tasks (Haiku 4.5, n=4); it reaches 94% where an agent over-builds (a date picker) and is near zero where the code is already minimal. ponytail keeps every safety guard while a bare “write one-liners” prompt drops one. (The earlier single-shot benchmark reported 80-94% as a flat figure; against a fair agentic baseline that is the per-task ceiling, not the average.) <a href=“benchmarks/results/2026-06-18-agentic.md”>Full writeup</a> · <a href=“benchmarks/“>reproduce it</a>.</sub> </p>

<p align=“center”> 代码减少约54%(最高94%)· 成本降低约20% · 速度提升约27% · 100%安全

<sub>实测数据来自Claude Code编辑真实开源项目(FastAPI + React)的会话,对比相同但无此技能的代理。54%是12个功能任务的平均值(Haiku 4.5, n=4);当代理过度构建(如日期选择器)时可达94%,在代码已最简处接近零。ponytail保留所有安全防护,而单纯”写单行代码”的提示会丢失一项。(早期单次基准测试报告的80-94%是统一数值;相对于合理的代理基准线,这是每项任务的上限而非平均值。) <a href=“benchmarks/results/2026-06-18-agentic.md”>完整报告</a> · <a href=“benchmarks/“>复现方法</a>.</sub> </p>

<p align=“center”> <sub><a href=“README.es.md”>Español</a> · <a href=“README.ko.md”>한국어</a></sub> </p>

<p align=“center”> <sub><a href=“README.es.md”>西班牙语</a> · <a href=“README.ko.md”>韩语</a></sub> </p>


<p align=“center”> <a href=“https://ponytail.dev/soon”&gt;![Something](assets/waitlist-banner.png)</a> </p>

<p align=“center”> <a href=“https://ponytail.dev/soon”&gt;![即将到来,加入等候名单](assets/waitlist-banner.png)</a> </p>

You know him. Long ponytail. Oval glasses. Has been at the company longer than the version control. You show him fifty lines; he looks at them, says nothing, and replaces them with one.

Ponytail puts him inside your AI agent.

你认识他。留着长马尾。椭圆眼镜。在公司的时间比版本控制系统还久。你给他看五十行代码;他看一眼,一言不发,然后换成一行。

马尾辫把他装进了你的AI代理里。

Before / after

改造前后

You ask for a date picker. Your agent installs flatpickr, writes a wrapper component, adds a stylesheet, and starts a discussion about timezones.

With ponytail:

你要一个日期选择器。你的代理会安装flatpickr、写封装组件、添加样式表,然后开始讨论时区问题。

使用马尾辫后:

 
<input type="date">
 
<input type="date">

More survivors in examples/.

更多幸存案例见examples/。

Numbers

数据

The honest measurement is a real agent doing real work: a headless Claude Code session editing tiangolo’s full-stack-fastapi-template (a real FastAPI + React repo), scored on the git diff it leaves behind. Twelve feature tickets, the same agent with and without the skill, n=4, Haiku 4.5.

真实测量来自实际工作场景:无头Claude Code会话编辑tiangolo的全栈fastapi模板(真实的FastAPI+React仓库),根据留下的git diff评分。12个功能需求单,同一代理启用/禁用该技能,n=4,Haiku 4.5。

<p align=“center”> Each arm as a percent of the no-skill baseline across LOC, tokens, cost and time (Haiku 4.5). ponytail is lowest on every metric (LOC 46%, tokens 78%, cost 80%, time 73%); caveman rises above 100% on tokens, cost and time; yagni-oneliner LOC 67%. Safety, separate adversarial tier: baseline, caveman and ponytail 100%, yagni-oneliner 95%. </p>

<p align=“center”> 各方案相对于无技能基准线的百分比(代码行数、token数、成本和时间,Haiku 4.5)。马尾辫在所有指标上最低(代码行数46%、token数78%、成本80%、时间73%);原始方案在token数、成本和时间上超过100%;YAGNI单行方案代码行数67%。安全性单独对抗测试:基准线、原始方案和ponytail 100%,YAGNI单行95%。 </p>

vs no-skill baselineLOCtokenscosttimesafe
ponytail-54%-22%-20%-27%100%
caveman+14%+37%+32%+18%100%
yagni-oneliner-33%---95%
对比无技能基准线代码行数token数成本时间安全性
马尾辫-54%-22%-20%-27%100%
原始方案+14%+37%+32%+18%100%
YAGNI单行-33%---95%

🔗 知识库双向关联