๐Ÿ›ก๏ธ AIcap Scan your repo free
2026-07-06 ยท AIcap Guides

Catching expensive GPU misconfigurations in Terraform and Kubernetes

AI infrastructure has a property regular infrastructure doesn't: the unit prices are violent. An p4d.24xlarge on-demand is roughly $24,000โ€“32,000 a month. A junior engineer's copy-pasted Terraform block can out-spend an entire team, and the feedback loop โ€” the cloud invoice โ€” arrives 30+ days after the merge.

The fix is the same one that works for security: move the check to the pull request. Your GPU footprint is declared in Terraform, Kubernetes, and Helm manifests, which means it is statically analysable in CI.

What a manifest scan can see

Each finding arrives with numbers attached, in the PR, where the conversation is cheap:

Resource: aws_instance.inference_node
Finding:  GPU instance g5.2xlarge without autoscaling policy
Cost:     $1.21โ€“1.45/hr โ†’ $880โ€“1,050/month (on-demand)
Spot:     $264โ€“315/month (~30% of on-demand)

The three findings that pay for the setup

1. Training-class hardware serving inference. The most expensive mistake in the catalog: p4d/p5-class instances (built for distributed training) running steady-state inference that an inf2 or g5 serves at a fraction of the price. Rightsizing detection compares the instance family against training signals in the repo (DVC pipelines, trainer imports) โ€” no training signals plus a training-class instance is a flag worth a human look.

2. On-demand where spot survives. Batch scoring, embedding backfills, and experimentation tolerate interruption; spot pricing runs around 30% of on-demand across the big three clouds. The scan projects both numbers so the PR shows the size of the decision, not just its existence.

3. Whole-GPU requests for fractional workloads. A K8s pod requesting nvidia.com/gpu: 1 for a model that uses 6 GB of an 80 GB A100 wastes most of a very expensive card. MIG partitioning or time-slicing fixes it; the scan flags the request that lacks either.

Why this lives in the compliance file too

The overlap nobody expects: EU AI Act Annex IV Section 2 asks for the system's computational resources and hardware requirements. If you are documenting a high-risk system anyway, the same manifest scan that guards your budget fills that section with real figures โ€” instance families, hourly ranges, monthly projections, and stated assumptions. One CI job, two audiences: the CFO and the auditor.

That is how the AIcap scanner packages it: FinOps findings render into Annex IV ยง 2(c) ("Hardware Requirements & Estimated Monthly Cost") with a total, a spot projection, and the assumptions block an auditor expects, while the same numbers show in the dashboard table for engineering review.

Honest limitations

Setup

It is the same scanner, the same job, zero extra flags:

- name: AI compliance + FinOps scan
  uses: istrategeorge/AIcap@v1.2.0
  with:
    scan-directory: '.'

Terraform, Kubernetes, and Helm files in the tree are analysed automatically; findings land in the AI-BOM JSON, the PR log, and the Annex IV draft. The first flagged p4d usually settles the question of whether the job was worth adding.

AIcap generates your AI-BOM, risk register, and Annex IV draft from a single CI run โ€” free CLI, EU-hosted ledger.

Get started free โ†’