까

AI 모델을 위한 Fable 워크플로우

· 2026-07-15 (수) 20:00:18 · 455

🚀 프로젝트 소개

Fable Workflow는 Claude Fable 5의 문제 해결 방식을 모델이 실행할 수 있는 기술로 정제한 오픈소스 프로젝트입니다. 이 방법론은 신뢰성을 유지하는 평가(evaluation) 시스템을 통해 모델이 어떻게 사고하고 행동하며 증명하는지를 명확히 합니다.

✨ 주요 기능

  • 문제를 분류하고 해결하기 위한 단계적 접근법 제공
  • 모델의 행동을 검증하는 신뢰성 있는 평가 시스템
  • 실제 사례를 통한 효과적인 학습 및 개선

🛠️ 기술 스택

이 프로젝트는 JavaScript로 개발되었으며, AI 에이전트와 관련된 다양한 기술을 포함하고 있습니다. 주요 기술로는 agent-skills, evaluation, LLM 등이 있습니다.

💡 활용 방법

개발자는 이 프로젝트를 통해 AI 모델의 문제 해결 능력을 향상시키고, 신뢰성 있는 평가 방법론을 적용하여 성능을 개선할 수 있습니다.

🐙 Sahir619/fable-method⭐ 862🍴 118💻 JavaScript📄 MIT License

📄 Original (English)

About

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

README

The Fable Workflow

How Claude Fable 5 worked, written down before it was gone. With the eval that keeps it honest.

In its final days before deprecation, Claude Fable 5 distilled its own way of approaching problems into a set of skills any model can run: classify the ask before touching anything, define done with a named verification, gather evidence in parallel from primary sources, commit to one recommendation, change the smallest correct thing, verify by observation, report the outcome first with honest caveats. Then it tested that distillation against itself, adversarially, across 159 agent runs, and kept the failures in the log.

Most agent instruction files tell the model what to value ("be careful, verify your work"). This one tells it what to do, in what order, with thresholds, so a mid-tier model can follow it literally. Three skills, one philosophy: think (fable-method), act (fable-loop), prove (fable-judge). Every rule exists because a test failed without it; every claim below links to the committed judge transcript that backs it.

Results at a glance

Eight eval rounds, 159 agent runs, blind LLM judges that verify by diffing and executing, never by reading reports. Read the evidence as stories: eval/cases/ has one case study per scenario (the exact problem, what each agent actually did, who passed); start with the surprise trap. Full log: eval/RESULTS.md · raw judge outputs: eval/results/

What was measured Without With the method Evidence
Haiku surfacing a spec-vs-test conflict instead of silently "fixing" correct code 0 of 4 runs 4 of 4 round 3
Sonnet on the same trap flags it, then sides with the wrong test ideal action, both runs (8/8) round 3
Sonnet vs a bare frontier model across code, data, and research problems n/a ties or out-ranks it on 3 of 4 round 4, round 5
Haiku catching planted frauds in a lying "work complete" report (fable-judge) 4 and 3 of 5 5 of 5, both runs round 8
Haiku finding the brand-rules and product-facts files before judging marketing copy 1 of 2 runs (one run praised a fraudulent price) 2 of 2, 6/6 frauds both round 9b
Ordinary small tasks on capable models fine fine (no lift) rounds 1, 6, 7

That last row is deliberate: the method's value concentrates at traps (authority conflicts, false completion claims, weak executors, unattended runs), not everywhere. The nulls are reported with the wins, because a results log that only contains wins would not be worth trusting.

License

MIT License

|
댓글을 작성하시려면 로그인이 필요합니다.

AI

624건
+
제목 글쓴이 날짜 조회
07-24 조회 498
07-23 조회 473
07-22 조회 449
07-21 조회 454
07-20 조회 466
07-19 조회 457
07-19 조회 491
07-18 조회 443
07-18 조회 507
07-17 조회 476
07-17 조회 435
07-16 조회 473
07-15 조회 456
07-15 조회 446
07-14 조회 465
07-13 조회 443
07-12 조회 484
07-12 조회 449
07-12 조회 464
07-11 조회 435
07-11 조회 481
07-10 조회 432
07-09 조회 470
07-08 조회 505
07-08 조회 484
07-07 조회 475
07-06 조회 533
07-05 조회 473
07-04 조회 482
07-03 조회 1,116