Six weeks with goose: an honest field report on a local AI coding agent

Three months with goose, an open-source AI coding agent, on my own hardware: a fast junior for the price of power. Great at bounded code, loops on vague work — and the harness matters.

by Attila Macskásy 6 min read

A single white goose standing calmly in a dark data-centre aisle between two long rows of server racks, blue status lights on the racks and one warm amber light falling on the goose

Everything you see on this blog is cooked in-house. The image above was generated on one of my own RTX 3090s, offline, in about eighty seconds — no stock photography, no cloud image API. Full recipe below.

Feature image made in my lab — image model, prompt and settings
Image model
Qwen-Image 2512 — fp8 (e4m3fn) variant. The bf16 pair does not fit a 24 GB card; fp8 weights are stored compressed and cast at compute time
Text encoder
Qwen2.5-VL 7B — a vision-language model doing the prompt understanding, which is most of why the composition follows a long prompt
Generated on
1× NVIDIA RTX 3090 (24 GB) in my own lab, through a self-hosted ComfyUI — my hardware, my electricity, nothing leaving the building
Settings
1664×928 · 20 steps · cfg 2.5 · euler / simple · shift 3.1 · seed 1995405906 · ~80 s
Prompt
a single white domestic goose standing calmly in the middle of a dark data centre aisle between two long rows of server racks, blue and cyan status lights glowing on the racks, one warm amber light falling on the goose, moody editorial photography, cold blue and cyan light with a single warm amber accent, shallow depth of field, 50mm, dark background, fine detail, cinematic, high dynamic range, no text
Negative prompt
text, words, letters, watermark, logo, signature, caption, ui, people, faces, hands, cartoon, illustration, cgi render, oversaturated, blurry, low quality, jpeg artifacts
Licence
Apache-2.0 — both the image model and the text encoder. Commercial use permitted with no use-based restrictions, which is exactly why this stack and not a prettier one

Update, 1 October 2026: I wrote this six weeks in; it goes out at roughly three months. goose is still installed, still wired to the same gateway, still on its own metered key. Alongside it now runs OpenCode, another open-source coding agent, also on its own key, so the two can be compared on real usage. Which one becomes my default local coding agent is still open — to be settled by measurement rather than argument, and there is no verdict yet. What still favours goose, as public facts: it runs natively on Windows, it is governed vendor-neutrally under the Linux Foundation (it moved there from Block, the company that started it), and it officially supports custom-branded distributions. Everything below is still how it went.

Why goose, and why local

I build and host a full platform on my own hardware. Frontier coding agents (Claude Code, Copilot) are excellent but proprietary: you cannot fork them, rebrand them, or run them air-gapped. So I went looking for an open, local-first agentic coder I could point at my own models on my own DGX cluster, with data never leaving the building. goose (Apache-2.0, Linux Foundation governance, native Windows CLI and desktop, a dedicated gateway provider) was the strongest candidate on paper, so I adopted it and put it to real work: building an actual product, phase by phase, against a metered LiteLLM gateway serving my coding model.

This is the honest report: the good, the bad, and the specific.

What goose is genuinely good at

  • Bounded backend code. Give it a well-scoped API phase — routers, queries, schemas, tests — and it delivers. Across several phases it produced ~2,400 lines of correct FastAPI/asyncpg code, following the existing module patterns closely.
  • Reading a codebase to act. It navigates, greps, reads the right files, and edits in place. The “junior developer who has read the repo” experience is real.
  • Running unattended. With auto-approve on file edits and a safe shell, it can grind through a phase while I do something else. For the right task that is a genuine multiplier at near-zero marginal cost: a local model has no per-token bill, just electricity.

Where it struggled, and the pattern behind it

  • It loops on reasoning-dense or ambiguous work. On a phase that needed non-trivial integration logic, it re-read the same files for ~2 hours and produced nothing committable, eventually hitting the context limit. That is the failure mode that hurts: not a wrong answer, but no answer after a long, hot, power-hungry grind.
  • It declares done before it is done. More than once it wrote a “completed / committed” status note and then did not finish the last step, the commit. Never trust an agent’s self-report of the final action; verify the repo state.
  • It cannot see. No agent driving a text-only code model can judge a rendered UI. Anything visual needs a human, or a different tool, to check.

The root cause of the looping matters, because it changes the fix. It is usually not “the model is too dumb”. It is tool-calling reliability: the agent needs the model to emit tool calls in the exact format its parser expects. goose is deliberately model-agnostic, and its docs say it works best with Claude-class models; on a non-Claude local model, malformed tool calls show up as exactly this kind of loop. (I later proved this by keeping the model fixed and swapping only the harness; the looping went away.)

The operating model that made it work

I stopped expecting the agent to do everything and adopted a senior + junior split:

  • Junior (goose, local model): pattern-heavy, well-scoped build phases. Cheap, unattended, on-prem.
  • Senior (a frontier model): the hard-reasoning phases, the architecture, and, critically, the visual/UX pass the junior is blind to, plus unblocking the junior when it loops.

With that split, goose earned its place: it did the volume work for the price of electricity, and the senior spent expensive tokens only where they mattered. The mistake is treating a local junior as a drop-in senior; framed as a junior, it is genuinely valuable.

Operational lessons (the unglamorous bits)

  • Heat and power are real. Sustained agentic sessions push a small GPU cluster hard. Watch package temperatures and room cooling; bursty coding with cool-downs between turns is fine, sustained soak is what to plan for.
  • Sessions accumulate context. One phase per session. Start fresh each phase, or you will hit the context ceiling mid-task.
  • The desktop app can wedge. An unclean close can leave zombie processes holding a single-instance lock, so the next launch silently quits. Kill the stragglers, clear the lock file, relaunch.
  • Keep a metered key per tool and per person. Even with a free local model, per-key metering tells you who, and what, burned which resources. It is invaluable when comparing tools.

The verdict

goose is a legitimate, capable, open, local agentic coder, and the right way to think about it is a fast junior that works for the price of power, paired with a senior for judgment and eyes. It got me from “is local agentic coding even viable?” to a working senior/junior pipeline shipping real features on my own hardware.

It also taught me the lesson that reframed everything: when a local coder loops, suspect the harness, not just the model. Acting on that took me somewhere better. That experiment (same model, different harness) is written up as its own piece; it is in the queue at nextpost.blog, and the votes decide when it comes.