Building Guardrails for AI Coding Agents in Go

Building Guardrails for AI Coding Agents in Go

How to turn recurring code review comments into automated checks, verify that they catch bugs, and decide what still needs human review. An AI coding agent can catch a bug in one review and miss it in the next. Adding instructions or another reviewer still leaves someone to read the code and decide whether it meets the requirement. Each review takes time and tokens, even when the same requirement has been checked before.I built a Go repository that turns selected code review requirements into automated checks. It combines linter configurations with behavioral tests, an architecture test that follows indirect imports, and scripts that validate the checks against deliberately broken code.I call these checks guardrails: conditions a code change must meet before we accept it. For example, a goroutine that must stop after cancellation gets a test for that behavior. A rule that report calculations must be independent of storage gets a check of the package dependency graph. The agent receives a diagnostic it can act on, then runs the same check after the fix.A passing test may miss a bug, and a failing command may indicate a build error or timeout. I validated the checks against both faulty implementations and failures of the tools themselves.The sections below explain how to construct these checks, establish what they detect, and identify the decisions that still need review. The examples use a model reporting service to make each requirement and failure reproducible.This article assumes you know Go tests, contexts, channels, and CI. The repository contains the source code, configurations, and reproduction commands. The command output shown here was obtained with Go 1.27.1 and golangci-lint 2.13.2. These versions are pinned so you can reproduce the results.How the checks fit together Each requirement needs a check that can observe the relevant failure. Naming rules can inspect source code. Cancellation tests must execute the goroutine and wait for it to finish. A restriction on indirect dependencies requires walking the package import graph. I combined these checks around the requirements of the reporting service: Area Requirement or review concern Check Formatting and names Follow agreed formatting and naming conventions gofmt and naming rules Complexity and duplication Flag code that exceeds agreed limits or repeats existing logic Static analysis, followed by review Errors and resources Close resources and return the errors required by the contract Linters and tests with controlled failures Concurrency Finish work after cancellation; detect leaks, deadlocks, and data races in exercised scenarios Synchronization tests, leak checks, and the race detector Architecture Keep calculations independent of storage, including through intermediate packages Import rules and a test that traverses dependencies Test quality Check the required results and detect selected behavior changes Result comparisons, field initialization rules, and mutation testing Verification scripts create faulty variants in temporary copies, run the relevant commands, and inspect their diagnostics. Separate cases check that build errors and timeouts are rejected as evidence of detecting a violation.These scripts validate the guardrails. The ordinary tests and linters remain the commands used to accept a code change. Keeping those purposes separate matters: a verification script can pass because it successfully caused a test to fail on broken code.Reproduce the checks. Run the commands from the root of the example repository. Select Go 1.27.1 for the terminal session first. In fish: set -gx GOTOOLCHAIN go1.27.1 go version golangci-lint version The first version check should report Go 1.27.1. The second should report golangci-lint 2.13.2 built with Go 1.27.1. The README covers tool and dependency installation.Set GOTOOLCHAIN again in each new session. The go 1.27.0 line in go.mod sets a minimum version; it does not pin the toolchain to 1.27.1.1. Automate formatting and naming checks. Formatting and naming establish the baseline for the rest of the checks. Once the team agrees on a convention, a tool can apply or check it on every change. gofmt applies standard Go formatting to files in the current directory and its subdirectories: gofmt -w . Naming conventions need separate rules. For example, the initialism ID should be capitalized. This declaration breaks that convention:type request struct { UserId int } The var-naming rule in revive reports the problem and the expected spelling:testdata/naming/naming.go:5:2: var-naming: struct field UserId should be UserID (revive) The agent can change UserId to UserID and run the check again. Choosing a name that captures the domain, such as account or customer, still requires a developer's judgment.Agree on a rule with the team and try it on the existing code before enabling it. Otherwise, a small task can accumulate unrelated renames.Configure and run the checks. To check formatting without changing files, use gofmt -l .. It lists files that need formatting, but still exits successfully. Configure CI to fail when that list is nonempty. The gofmt documentation describes the flags. I run revive through golangci-lint, which runs multiple analyzers. The linter catalog lists the available checks. To verify the naming rule in isolation, this configuration enables it without unrelated rules. Save it as naming.yml:version: "2" linters: default: none enable: - revive settings: revive: enable-default-rules: false rules: - name: var-naming This disables two sets of defaults. default: none disables golangci-lint's default linters; enable-default-rules: false disables revive's default rules. The resulting diagnostic can then be checked against the specific naming violation. The configuration documentation explains the structure.Run this from the example module directory:golangci-lint run --config naming.yml ./testdata/naming The path is explicit because ./... skips testdata. The command reports the diagnostic above, including the file, location, and linter name.2. Define what complexity checks should enforce. Working code can still become harder to maintain after an agent's change. Nested conditions add paths to follow, and duplicated logic creates multiple places to update. Static analysis can flag these changes for review, but the rule must express what the team wants to control. A complexity threshold sets an upper limit for a function. Preventing complexity from increasing requires a comparison with the previous version. These are different acceptance criteria: with a threshold of 30, a score rising from 20 to 25 still passes.Make the threshold violation reproducible. For the repository's complexity check, I used gocognit and a function whose nested conditions have a predictable score. CanExport allows an export when all three conditions hold: func CanExport(active, allowed, ready bool) bool { if active { if allowed { if ready { return true } } } return false } gocognit assigns a cognitive complexity score to control flow. Here, the three nested if statements contribute 1, 2, and 3, for a total of 6. Each level of nesting increases the contribution. The gocognit rules explain the calculation.I set the threshold to 3 so this implementation triggers a known diagnostic:cognitive complexity 6 of func `CanExport` is high (> 3) This threshold is part of the test case. Choose a threshold for a project after examining its existing code and agreeing on the limit with the team.Combining the conditions preserves the behavior:func CanExport(active, allowed, ready bool) bool { return active && allowed && ready } The diagnostic identifies the function for the agent to change. Review must still assess the resulting code: splitting a function into many helpers can lower its score while making the execution flow harder to follow.Keep the finding visible in the diff. An existing project may have many violations when a rule is first enabled. Filtering findings to changed lines helps focus on a patch, but the location of a complexity diagnostic can make that filter hide a new violation. The diagnostic may point to the function declaration. If the agent changes only the body, the declaration is outside the diff. The linter can detect the violation, and the filter can remove it from the reported results.With revision filtering, --whole-files reports findings for the entire changed file. That includes the unchanged declaration, but also brings back existing violations elsewhere in the file. The CLI documentation explains these filters.If the requirement is to prevent any increase in complexity, compare the function's score before and after the change, using a branch such as main as the baseline. The threshold check shown here does not implement that comparison. Choose the acceptance criterion and reporting scope together so that a relevant finding reaches the agent and reviewer.Use other metrics to focus review. Complexity is one source of maintenance work. Other analyzers identify code that needs a different review decision: Tool What it finds funlen Functions that exceed configured length limits dupl Similar code fragments that may be duplicates unused Unused declarations in the analyzed packages A duplication finding is a prompt to decide whether the fragments should share an implementation. Introducing an interface or wrapper still needs a design justification. Analyzers also have technical limits: unused, for example, does not flag an exported declaration solely because nothing in the project calls it.3. Verify resource cleanup and error propagation. A resource check needs to cover what happens when an operation fails. For Download, I defined the contract after a successful HTTP request: close the response body exactly once, preserve any data already read, and return both the read and close errors when both occur. I paired a linter check for the missing close with a behavioral test of that contract. The test controls read and close failures independently, including the case where reading returns data before failing.Close the body and preserve both errors. This implementation reads the response body but leaves it open: func Download(url string) ([]byte, error) { response, err := http.Get(url) if err != nil { return nil, err } return io.ReadAll(response.Body) } bodyclose detects the missing close in this implementation. Closing the body through defer handles cleanup when reading succeeds or fails. The deferred function also needs to preserve any close error required by the contract:func Download(url string) (data []byte, err error) { response, err := http.Get(url) if err != nil { return nil, err } defer func() { err = errors.Join(err, response.Body.Close()) }() return io.ReadAll(response.Body) } The return statement assigns the results of io.ReadAll to the named return variables, data and err. The deferred function then closes the body and uses errors.Join to combine any close error with the read error. The caller receives the result after that deferred function finishes.When both operations succeed, the returned error is nil. When either fails, the caller can inspect the returned error with errors.Is. The function may return data together with an error, so receiving data alone does not establish success.Test the failure combinations. I used TestDownloadReadAndClose to check all four combinations of read and close outcomes: Read outcome Close outcome Expected error Success Success nil Failure after returning data Success Contains the read error Success Failure Contains the close error Failure after returning data Failure Contains both errors Every case also checks that Close is called exactly once and that the returned data matches what the body supplied. Checking partial data matters: a change that discards those bytes would violate the contract even if it closes the body and returns the expected error.The test supplies an http.RoundTripper that returns a controlled response without making a network request. Its body counts calls to Close and can return a configured close error. For read failures, the reader supplies data first, then returns the configured error. This makes each failure combination reproducible without depending on network behavior.Because Download calls http.Get, the test temporarily replaces http.DefaultClient. The subtests run sequentially and restore the original client after each case. Run the test with:go test -race -count=1 -run '^TestDownloadReadAndClose$' ./internal/download After the fix, the bodyclose finding disappears and the function passes the repository's baseline linters. The behavioral test checks the contract in the four scenarios above. bodyclose recognizes known code patterns, and passing these checks does not prove cleanup on every possible execution path.The scope here is response body cleanup and error propagation. Timeouts, context propagation, and HTTP status handling need separate requirements and checks.Apply cleanup checks to other resources. A context with a timeout also needs cleanup. context.WithTimeout returns a cancel function; deferring it releases the associated resources when the function returns, without waiting for the timeout: ctx, cancel := context.WithTimeout(parent, timeout) defer cancel() Replacing cancel with _ causes the lostcancel check in go vet to report the missing cancellation call.The errcheck linter helps find ignored errors. If code checks an error and then returns success, nilerr detects some of those cases.For each dependency, define what its failure should mean to the caller. Then make the dependency fail in a test and check the required response. That contract determines whether the function should propagate an error, preserve partial results, or return another agreed outcome.4. Test cancellation, deadlocks, and data races. For concurrent code, I built separate checks for completion after cancellation, blocked goroutines, and conflicting access to shared data. Each check needs a scenario that reaches the failure and a diagnostic that identifies it. A timeout alone cannot distinguish a deadlock from slow execution. The cancellation test starts with an observable contract: while no receiver is available, the sender waits; after cancellation, it finishes and signals completion.Make completion observable. A goroutine sends a result on an unbuffered channel. If the receiver stops reading, the send stays blocked. Handling context cancellation gives the goroutine a way to stop waiting. Forward sends a value to out and closes done when it finishes. The caller can wait on done for completion:func Forward(ctx context.Context, out chan internal/shared -> internal/storage The architecture rule covers the entire import chain, including dependencies through shared packages. The configured depguard rule accepts this change because calculation code does not import storage directly. But the dependency exists through shared.The architecture test loads the packages, confirms that the required packages exist, and walks their imports. If calculation code can reach storage, it fails and reports the dependency path:ARCH001: calculation depends on storage: example.com/guardrails/internal/reportcalc -> example.com/guardrails/internal/shared -> example.com/guardrails/internal/storage The agent can see which package introduced the forbidden dependency. Fixing it may require moving a shared calculation or changing the dependency direction while preserving program behavior.I also made configuration failures visible. After a package rename, an old path could stop matching anything and leave the rule checking no code. The test fails if either side of the rule is missing or the package graph cannot be loaded. A successful result therefore requires the intended packages to be present and checked.Its scope is the import graph for the selected build. If platforms or build tags change the files included, check those configurations separately. Accessing storage over the network can happen without a Go package dependency, so this test will not catch it.Load and validate the package graph. The test loads packages with golang.org/x/tools/go/packages. NeedDeps supplies dependencies for traversal. I also requested type information and checked loading errors so an incomplete or invalid graph cannot produce a passing result. In this code, calculationRoot is example.com/guardrails/internal/reportcalc, and storageRoot is example.com/guardrails/internal/storage. Tests: false excludes test files, matching the depguard configuration above. The helpers and their tests are in architecture_test.go.func TestArchitecture(t *testing.T) { pkgs, err := packages.Load(&packages.Config{ Mode: packages.NeedName | packages.NeedImports | packages.NeedDeps | packages.NeedTypes, Tests: false, }, "./internal/...") if err != nil { t.Fatal(err) } if len(pkgs) == 0 || packages.PrintErrors(pkgs) > 0 { t.Fatal("cannot load the production package graph") } targetFound := false for _, pkg := range pkgs { if pkg.PkgPath == storageRoot { targetFound = true } } if !targetFound { t.Fatal("architecture target package missing; review storageRoot") } checked := 0 for _, pkg := range pkgs { if !inside(pkg.PkgPath, calculationRoot) { continue } checked++ if chain := forbiddenPath(pkg, storageRoot, make(map[string]bool)); chain != nil { t.Errorf("ARCH001: calculation depends on storage: %s", strings.Join(chain, " -> ")) } } if checked == 0 { t.Fatal("no calculation packages checked; review the architecture rule") } } inside selects a package and its subpackages, respecting the / separator. A rule for storage therefore does not include storagecache.forbiddenPath walks imports and returns a chain to the forbidden package. It skips packages already visited and sorts imports before traversal. When several forbidden paths exist, the diagnostic consistently reports the same one.I validated these safeguards with separate cases for a missing storage package, an empty set of calculation packages, and a graph that fails to load. The package boundary cases check a direct dependency, a dependency on storage/reader, and the allowed neighboring package storagecache. These cases test both missed violations and incorrect rejections.Build tags are easy to lose. go test -tags does not automatically pass its tags to the nested packages.Load call. The test can run with tags while loading a graph without them. Set them through BuildFlags or a shared GOFLAGS environment variable; the repository checks the latter. Test files also need explicit inclusion if you want to inspect them. See the go/packages configuration.6. Check whether tests catch bugs. The agent's tests become part of the acceptance criteria, so they need scrutiny too. I included faulty implementations that expose two ways a test can pass incorrectly: deriving its expectation from the same faulty expression and comparing only part of the result. Mutation testing then checks whether selected changes in behavior cause the tests to fail. Derive expectations from the requirement. When an agent writes both the implementation and the test, it can make the same mistake twice. Suppose the requirement allows values up to and including 10, but the function rejects 10: func Allowed(size int) bool { return size ) /example/internal/report/report_test.go:33 + goroutine 8 [chan send (durable), synctest bubble 1]: example.com/guardrails/internal/report.Forward.func1() /example/internal/report/report.go:26 + The test goroutine is waiting to receive from done; the sender is blocked at out <- value. While the send is blocked, defer close(done) cannot run. The stack identifies both sides of the wait.Fix the cause and rerun the command Add the cancellation case to select: select { case out <- value: +case <-ctx.Done(): } This restores the implementation from section 4. Run the same command again:go test -count=1 -timeout=10s -run '^TestForwardCancellation$' ./internal/report After the fix, it exits with code 0. One run produced:ok example.com/guardrails/internal/report 0.258s Execution time depends on the machine.Next, check that the fix preserved the other behavior. Run the whole module's tests, including successful delivery:go test -race -count=1 -timeout=30s ./... That run passed too. In the proposed workflow, CI repeats this command on the revision under review using the same Go version. The cancellation test is part of the full suite.Validate a guardrail before making it mandatory. The repository's verification scripts check both the code rules and the interpretation of their results. The same validation process can be applied before making a new check mandatory in a project: validate the configuration, test a known violation, and inspect findings on the existing code. Validate the rule itself. Start with configuration validation. For the complexity example: golangci-lint config verify --config complexity.yml The command validates settings against a schema and catches typos such as min-complexty, which an ordinary run may ignore. Valid settings can still target the wrong code. Run the rule against code that should pass and a known violation that should produce the expected diagnostic.Check the message as well as the exit code. A tool that failed to start can also return a nonzero status.I implemented this validation in verify.py and verify_mutation.py. They run faulty variants in temporary copies and inspect structured output. The Go test validator checks that the target package and test ran, then checks the expected diagnostic. The linter validator checks the analyzer name, source position, and diagnostic text.The saved verification.json report records 50 successful verification steps on Go 1.27.1 and golangci-lint 2.13.2, running on macOS arm64. Those steps include configuration validation, baseline checks, known violations, and cases that must be rejected as invalid evidence. The separate mutation-verification.json report records the weak and boundary test suites and the four invalid-result controls.For these scripts, detecting an expected test failure is a successful verification step. Acceptance of a code change still requires the ordinary tests and linters to pass.Pay particular attention to settings that select files, packages, or types. A valid but incorrect path can leave the intended code unchecked. The storageRoot path and the View regular expression above are examples. The README covers offline configuration validation and reproduction commands.Review findings before blocking changes Run the rule on the existing code. Fix configuration errors, agree on exceptions, and review the reported violations before making it block changes. If there are many violations, first collect a report, then prevent new ones while fixing the old ones gradually. Remember that filtering by changed lines can hide new findings, as the complexity example showed.Connect the agent and CI Give the agent the check command. Its diagnostics should identify the cause: A linter should name the rule, file, and line. A test should identify the scenario, expected result, and observed result. An architecture test should show the forbidden dependency chain. The workflow is a short loop: change the code, run the check, inspect the diagnostic, fix the cause, and rerun the same check. Once the checks pass, review the change.Run quick checks after edits and schedule slower tests and mutation runs at selected stages. Checks used to accept a change must also run in CI on the revision being reviewed.If a tool cannot run, fix the execution problem and rerun it. A tooling failure tells you nothing about whether the code is correct.The feedback loop distinguishes rule violations, passing checks, and check execution failures. Keep code review focused on the contract. The repository combines executable requirements with checks of their scope and diagnostics. It covers behavior such as cancellation and resource cleanup, architecture rules that follow indirect dependencies, and tests of whether the guardrails detect known violations. Review still needs to establish whether the contract is right for the application. That includes deciding whether a send is allowed when cancellation is also ready and whether report calculations should depend on storage.Check the scope too. The cancellation test for Forward does not cover a new goroutine in another function. Linters and architecture tests cover only the code their configuration selects.Changes to the checks deserve the same attention as changes to the code. Raising a threshold, adding an exclusion, deleting an assertion, or splitting a function just to lower its score can remove a finding while leaving the original problem.The recorded runs establish that the selected checks detect the intended violations and that the validators reject the tested forms of invalid evidence. Agent behavior and token savings were not measured.To apply this approach, choose a recurring review requirement and write down its acceptance criterion. Keep a known violation alongside the check, verify its diagnostic, and give the agent a command it can rerun after a fix. Review changes to that check against the same requirement.

Original Source

Read the full article at Hackernoon →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.