Agentic engineering · Aotearoa New Zealand

Build software with AI agents — properly

From clever demos to dependable delivery: agents that plan, code, test and ship, with engineers firmly in charge.

Agentic engineering is the discipline of putting AI agents to work across the software lifecycle — not just autocomplete, but agents that take a ticket, explore a codebase, make changes, run the tests and open a pull request. Done well, it lets small Kiwi teams punch well above their weight. Done carelessly, it ships confident mistakes faster. This site is about doing it well.

What is agentic engineering?

It's the shift from prompting a chatbot for snippets to designing systems where agents do real engineering work inside clear boundaries.

The model is only one part. The engineering is in the context you give it, the tools it can use, the checks it must pass, and the moments a human signs off.

  • Agentic coding — agents that read, plan, edit and refactor across a real repository, not a blank chat window
  • Autonomous dev workflows — background agents that pick up well-scoped tickets, triage CI failures and open reviewable PRs
  • Evals and testing — measuring agent output against tests and task suites, so “it looks right” becomes “it passes”
  • Orchestration — planners, sub-agents and tool calls coordinated so long tasks stay on track
  • Guardrails and human-in-the-loop — sandboxes, least-privilege credentials and approval gates where they matter

The practices that make it work

Six building blocks we keep coming back to when teams move from experimenting with agents to relying on them.

Agentic coding

Agents working in your actual codebase, with your conventions.

  • Repo-level instructions (e.g. an AGENTS.md)
  • Small, reviewable diffs over big-bang rewrites
  • Plan first, then edit, then verify

Autonomous workflows

Agents running in the background on well-defined jobs.

  • Ticket → branch → pull request
  • CI failure triage and flaky-test fixes
  • Dependency bumps and routine upkeep

Evals & testing

Trust comes from measurement, not vibes.

  • Tests as the agent's definition of done
  • Task suites to compare models and prompts
  • Regression checks when anything changes

Orchestration

Coordinating agents, tools and context over longer tasks.

  • Planner and sub-agent patterns
  • Tool access via standards like MCP
  • Context management and hand-offs

Guardrails

Safe by design, not by hope.

  • Sandboxed execution environments
  • Least-privilege, scoped credentials
  • Prompt-injection awareness and audit logs

Human-in-the-loop

Engineers stay accountable for what ships.

  • Approval gates for irreversible actions
  • Code review that's built for agent PRs
  • Clear ownership when things go wrong

Example: an agent task, written like an engineer would

A good agent brief reads a lot like a good ticket: clear scope, the tools allowed, how “done” is checked, and where a human must approve. This illustrative spec isn't tied to any one product.

# agent-task.yaml — illustrative only
task: "Add NZ GST (15%) breakdown to invoice PDF"
repo: billing-service
context:
  - AGENTS.md
  - docs/invoicing.md
tools:
  allowed: [read_files, edit_files, run_tests, open_pull_request]
  denied:  [deploy, prod_database, send_email]
sandbox: true
definition_of_done:
  - "npm test passes"
  - "new unit tests cover GST rounding"
  - "eval suite: invoices/* no regressions"
human_in_the_loop:
  review_required: true
  approve_before: [merge, release]
budget:
  max_steps: 40
  on_limit: "stop and summarise progress"

The agentic loop

Most dependable agent workflows follow the same shape. The craft is in tightening each step.

  1. 1 · Scope

    Frame the task

    Write a brief with a clear goal, constraints, relevant context and an explicit definition of done.

  2. 2 · Plan

    Let the agent plan, then check it

    The agent explores the code and proposes an approach. A quick human glance here saves a lot of rework later.

  3. 3 · Build

    Act inside a sandbox

    Edits, commands and tool calls happen in an isolated environment with only the permissions the task needs.

  4. 4 · Verify

    Prove it with tests and evals

    Automated checks decide whether the work is done — and failures feed straight back into the next attempt.

  5. 5 · Review

    A human signs off

    Engineers review the diff and the reasoning, then merge. Anything irreversible waits for explicit approval.

Who it's for

Any New Zealand team that builds or maintains software and wants AI agents to be a genuine part of how the work gets done.

Startups & scale-ups

Small teams that need to ship more without burning out

Enterprise engineering

Platform and delivery teams standardising agent use safely

Public sector

Agencies balancing productivity with privacy and assurance

Software agencies

Studios and consultancies delivering client work with agents

Platform & DevOps

Teams building the sandboxes, pipelines and guardrails agents run in

Tech leaders

CTOs and heads of engineering shaping how their teams adopt agents

If your team writes software, agents are about to become part of the team.

Engineering first. Agents second.

The teams getting real value from AI agents aren't the ones with the flashiest demos — they're the ones with good tests, clear briefs, tight permissions and a human who owns the outcome. Start there, and the agents get a lot more useful.

Explore the AXM network