Spots

GitHub - agustindiazcano/afaw-anti-fragile-agentic-workflow: AFAW: AI-assisted…

A deterministic, multi-agent methodology for AI-assisted software development.

Several coding agents working on one repository fail

Several coding agents working on one repository fail in predictable ways: they collide on branches and on shared context files, claim results they never measured, and write tests that pass whatever the code does. AFAW is a repository-level method and boilerplate that lets agents write code in parallel while deterministic tools decide whether the result is acceptable and a human approves every merge. AFAW is the method; the code here is one way to implement it. See what is core and what you can replace.

A dashboard built by scripts/build_state.py from mock data

A dashboard built by scripts/build_state.py from mock data: a fictional project whose every task and number is invented (examples/mock_dashboard.py). It shows each traffic-light rule firing: CI failing, a blocker not done, no activity for three days, nothing measured. Build it locally: python -m examples.mock_dashboard mock-output and open mock-output/index.html.

It costs less. Without it, an agent runs

It costs less. Without it, an agent runs the whole test suite on your machine, the machine is slow, tests fail for reasons that have nothing to do with the code, and the agent retries for an hour. Every retry sends its whole context again: that is where the token bill goes. With AFAW the agent runs only the test it is writing; the cloud runs the rest in parallel and answers once. A few minutes of CI cost far less than an hour of an agent going in circles.

It works alone and in a team. Alone

It works alone and in a team. Alone, you let the agents run and nobody checks them: they commit to the wrong branch, write tests that pass no matter what, and report results they never measured. The checks, the branch protection and the hooks do that reviewing for you.

With 2 or 10 people, everyone sees what

With 2 or 10 people, everyone sees what the agents did. Each person runs their own agents, and without a shared record nobody knows what the others' agents changed, why, or where each task stands. Here every task, decision and status is in the same place and in the same format, for every person and every agent: the board, the generated LASTCONTEXT.md and the decision records. You set it up once, from the template.

The project documents itself. Every task leaves what

The project documents itself. Every task leaves what was done, what is left and, for each decision, why it was taken. When agents do the work, that is usually lost: afterwards nobody knows what was done or why.

A new agent understands the project in minutes

A new agent understands the project in minutes. Instead of reading every file from scratch, it reads one generated index (LASTCONTEXT.md) and follows links to the decisions it needs. Less reading means fewer tokens and less time, and the saving grows with the project.

Fewer status meetings. "What did you do, where

Fewer status meetings. "What did you do, where does it stand, what's blocked, why was this decided": the board and the records already answer it, and an optional assistant can answer it in a chat. See Ask the project.

These are the expected gains, not measured ones

These are the expected gains, not measured ones yet. The dashboard already records timings and CI cost, and the paper's evaluation protocol says how to measure all four. Method and this implementation

News

GitHub - agustindiazcano/afaw-anti-fragile-agentic-workflow: AFAW: AI-assisted development framework for deterministic multi-agent orchestration.

A deterministic, multi-agent methodology for AI-assisted software development.

@spots #dev
Source: Show HN
See more like this