FinOps Maturity Assessment

Generated 2026-07-23T07:04:42.250Z · Engine finops-1.0.0

Model routing mode: normal

Models: gpt-5.5 (Pre-Flight) · claude-sonnet-4-6 (Audit) · gpt-5.5 (Evidence Check) · claude-sonnet-4-6 (Summary/Diagnosis) · claude-opus-4-7 (Roadmap) · gpt-5.5 (Fact-Check)

Knowledge Base: Remote PDF KB loaded (60 PDFs)

Source parse note: Source packet F has weak deterministic routing coverage (4/4 chunks); broad-source fallback was used for that batch.

Evidence Check

Quality Gate Status: WARN

Assessment score remains valid. Unsupported strategy wording or actions were removed or retained only in the appendix.

17/31 claims supported

Phase 1 findings were verified against the raw material before Phase 2 metrics were calculated.

Supported49
Weak9
Unsupported2
Missing0
Downgraded22
Rescanned20
Adjusted criteria
maturity.A2 · 2→1 · supported · rescanned
The source supports a functioning weekly showback model using tags, account hierarchy, and shared platform allocation rules. The scanner correctly withheld credit for formal stakeholder acceptance and regular refinement of the allocation model, which are not evidenced.
maturity.A4 · 3→2 · supported · rescanned
The dashboard evidence supports role-appropriate executive, finance, and engineering views, 12-hour refresh, trend/forecast information, and AWS/Azure/GCP billing export integration. The scanner correctly withheld full credit because SaaS platform integration is not evidenced.
antipattern.A1 · 1→0 · supported · rescanned
The raw material shows the opposite of tag sprawl: mandatory tag keys, daily defect reporting, remediation SLA, exception expiry, and policy-as-code checks for new Terraform modules. No evidence supports significant untagged spend, ad-hoc taxonomy, or acceptance of untagged spend as overhead.
antipattern.A5 · 1→0 · unsupported
Adjudication: The cited sources show limited unit economics maturity and some savings reported as absolute EUR run-rate reduction, but they do not show misleading, unverifiable, presentation-only, or non-actionable vanity metrics. The optimization evidence indicates operational use through right-sizing recommendations, product-team approval/rejection, and backlog decisions.
maturity.B1 · 3→2 · supported · rescanned
The cited chunks support active RI/Savings Plan/CUD management across AWS/Azure/GCP, a defined 72% coverage target, current 68% coverage tracking, monthly utilization review, escalation of unused commitments, dashboard utilization tracking, and quarterly commitment strategy review. The source does not show exchange/resale/remediation of unused commitments, so a score of 2 is appropriate.
maturity.B2 · 3→2 · supported · rescanned
The source supports right-sizing recommendations generated weekly from CPU, memory, and storage utilization, product-team review in an optimization backlog, weekly backlog review, and tracked financial impact of EUR 46,000 monthly run-rate reduction. Automatic generation is not explicit, so criterion 2 is only partially supported; count 2 is reasonable.
maturity.B4 · 2→1 · supported · rescanned
The source supports spot/preemptible approval for batch analytics, CI test runners, and image-processing jobs. It also explicitly limits maturity: interruption handling exists only for analytics batch jobs, CI fallback remains manual, and expansion targets are not approved. Count 1 is supported.
antipattern.B2 · 1→0 · unsupported
Adjudication: The only harmful signal is that 7 right-sizing recommendations were deferred due to release freeze risk, while 22 were accepted and reduced monthly run-rate by EUR 46,000. The source does not show chronic over-provisioning, utilization below 20%, unvalidated peak-capacity justification, or ignored recommendations due to lack of ownership.
antipattern.B4 · 1→1 · weak
The cited evidence shows some manual or incomplete automation gaps: CI runner fallback is manual, Infracost is not enforced across all product repositories, and legacy manually provisioned resources remain. However, these do not support the Manual-Only Optimization anti-pattern criteria. The source also documents weekly generated optimization recommendations, weekly backlog review, automated anomaly detection, dashboarding, and non-production shutdowns. The positive score is too strong for the harmful pattern.
antipattern.B5 · 0→0 · weak
The source supports commitment utilization tracking and multi-cloud commitment management, which argues against a total absence of rate optimization governance. However, it does not specifically evidence discount-interaction understanding, overlap analysis, or avoidance of double-coverage. The scanner's zero positive score is acceptable, but its claim that the source directly contradicts all discount-stacking ignorance criteria is overstated.
maturity.C1 · 3→2 · supported · rescanned
The source supports a documented governance framework, tagging requirements, exception approval, named governance ownership, and some policy-as-code enforcement in Terraform modules. It also supports the scanner's caveats: spend limits/full provisioning approval workflows and policy lifecycle review cadence are not explicit.
maturity.C2 · 3→2 · supported · rescanned
src-001-c001 directly supports monthly product-team budgets, monthly budget-to-actual variance checks above 10%, a forecasting model using historical spend/planned launches/commitment expiries/traffic growth, and corrective actions in a monthly FinOps action log. The scanner appropriately did not award full credit because project/org-level budget hierarchy is not explicit.
Walk | Delivery 100% · Evidence 57%
Evidence-Gated Readiness
50%
Evidence density 57% is below 60%, so readiness is capped until more current-state evidence is supplied.
Maturity Depth
47%
Average maturity score across all criteria on a 0–3 scale, normalized to 0–100%. Captures partial progress that maturity_ratio misses.
Anti-Pattern Burden
7%
Average severity across all anti-patterns. Higher = more friction blocking current FinOps practice. Low values mean "low confirmed burden" only when source evidence is strong enough.
Anti-Pattern Clearance
63%
Share of anti-patterns that were meaningfully tested and not found. This is positive only when the source had relevant coverage.
Anti-Pattern Coverage
83%
Share of anti-pattern criteria that were meaningfully assessed, either as findings or verified absences. Low coverage means absence is unknown, not good.
Maturity Ratio
17%
Share of maturity criteria that scored as fully embedded (3 of 3 sub-criteria met).

Maturity Gauges

50%

Evidence-Gated Readiness

Evidence density 57% is below 60%, so readiness is capped until more current-state evidence is supplied.

Target: High
17%

Maturity Level

Share of maturity criteria that scored as fully embedded (3 of 3 sub-criteria met).

Target: High
47%

Maturity Depth

Average maturity score across all criteria on a 0–3 scale, normalized to 0–100%. Captures partial progress that maturity_ratio misses.

Target: High
20%

Anti-Pattern Level

Share of anti-patterns scored as deeply entrenched (3 of 3 sub-criteria met). Higher = worse.

Target: Low
7%

Anti-Pattern Burden

Average severity across all anti-patterns. Higher = more friction blocking current FinOps practice. Low values mean "low confirmed burden" only when source evidence is strong enough.

Target: Low
63%

Anti-Pattern Clearance

Share of anti-patterns that were meaningfully tested and not found. This is positive only when the source had relevant coverage.

Target: High
83%

Anti-Pattern Coverage

Share of anti-pattern criteria that were meaningfully assessed, either as findings or verified absences. Low coverage means absence is unknown, not good.

Target: High
100%

Delivery Integrity

Did the audit pipeline complete? Share of maturity and anti-pattern criteria the LLM returned valid data for. Below 100% means batches failed.

Target: High
57%

Evidence Density

Did the source actually cover the criterion? Share of maturity and anti-pattern criteria with verified source coverage, including positive evidence, quote-backed gaps, anti-pattern findings, and verified anti-pattern absences.

Target: High

Visual Diagnosis

Category Footprint

Per-domain maturity (emerald) vs anti-pattern burden (rose). Each axis is one assessment domain; values are the sum of sub-criterion counts (0–15) for that domain.

A · VisibilityB · OptimizationC · GovernanceD · ArchitectureE · CultureF · GenAI Maturity Anti-Patterns

Position vs. Quadrants

Validated maturity depth (x-axis) plotted against confirmed anti-pattern burden (y-axis). When evidence or anti-pattern coverage is insufficient, quadrant labels are suppressed.

COST BLINDNESS FINOPS THEATER LOW / UNPROVEN SIGNAL VALIDATED STRENGTH Validated Maturity → Confirmed Burden → 0% 100% 100% 0% 47% / 7%

Domain Signal Overview

Maturity target is high; anti-pattern finding rate target is low. Grey means the source did not provide enough assessable coverage.

A

Cost Visibility & Allocation

Maturity signal
60%

5/5 criteria assessed

Anti-pattern finding rate
0%

0 findings, 0 partial, 1 not assessed

B

Rate & Usage Optimization

Maturity signal
53%

5/5 criteria assessed

Anti-pattern finding rate
20%

0 findings, 2 partial, 1 not assessed

C

Governance & Policy

Maturity signal
60%

5/5 criteria assessed

Anti-pattern finding rate
10%

0 findings, 1 partial, 0 not assessed

D

Architecture & Engineering

Maturity signal
53%

4/5 criteria assessed

Anti-pattern finding rate
30%

0 findings, 3 partial, 0 not assessed

E

Culture & Organization

Maturity signal
53%

4/5 criteria assessed

Anti-pattern finding rate
0%

0 findings, 0 partial, 1 not assessed

F

GenAI & AI Cost Management

Coverage note
Maturity signal
0%

0/5 criteria assessed

Anti-pattern finding rate
0%

0 findings, 0 partial, 2 not assessed

Anti-pattern absence is not fully assessable from source coverage.

Evidence Summary

Fact-only current state · Walk

Walk-maturity cloud FinOps program with documented governance and multi-cloud visibility, constrained by enforcement gaps, a 57% evidence density ceiling, and a complete absence of GenAI cost management evidence.

Key metrics

  • Evidence-Gated Readiness Score: 50/100 (capped; evidence density 57% below threshold)
  • Maturity Depth Index: 47%
  • Anti-Pattern Burden: 7% (confirmed, 6 anti-patterns)
  • Anti-Pattern Clearance: 63%
  • Anti-Pattern Coverage: 83%
  • Delivery Integrity: 100%
  • Evidence Density: 57%
  • Maturity Gaps: 7
  • Silent Areas: 7
  • Domain scores — A: 9/15, B: 8/15, C: 9/15, D: 8/15, E: 8/15, F: 0/15
  • Commitment coverage: 68% actual vs. 72% target (stable compute and database)
  • Idle resource waste: 4.8% of monthly spend vs. below-3% 2026 target
  • Right-sizing accepted in 2026-Q1: 22 recommendations; EUR 46,000 monthly run-rate reduction
  • Anomaly event (2026-03-18): EUR 18,400 estimated monthly avoided cost

Confirmed strengths

  • Multi-cloud FinOps dashboard (AWS, Azure, GCP) refreshing every 12 hours with drill-down to service and owner_team level
  • Mandatory tagging policy covering six tag keys (business_unit, product, environment, owner_team, cost_center, data_classification) with daily defect reporting and five-business-day remediation SLA
  • Weekly cost allocation showback to product teams using tags, account hierarchy, and shared platform rules
  • Automated daily anomaly detection with alert routing to owning team Slack channels and one-business-day acknowledgement requirement for high-severity alerts
  • Federated FinOps operating model with documented RACI across Finance, FinOps, Platform Engineering, and product teams
  • FinOps Council with cross-functional representation meeting monthly and reporting to CTO/CFO steering group
  • Quarterly commitment and architecture cost reviews with CTO/CFO steering group documented
  • Architecture review templates requiring estimated monthly cost, scaling cost curve, data-retention cost impact, and cost-performance tradeoff section
  • Autoscaling policies required to define maximum capacity, target utilization, and cost ceiling guardrails
  • Terraform standard for production infrastructure with mandatory tags and policy-as-code checks in new modules
  • Infracost estimates visible in pull requests for the platform account
  • Non-production environment shutdown policies with always_on exception approval process
  • Weekly waste detection for idle load balancers, unattached disks, and orphaned snapshots
  • Engineering cost owner [PERSON_NAME_REDACTED] product domain
  • Weekly optimization backlog review cadence with product cost owners documented

Confirmed gaps

  • Chargeback not yet active; showback-only allocation in place (chargeback pilot planned for 2026-Q3 for shared analytics workloads)
  • Unit economics defined only for checkout transactions and loyalty API calls; cost-per-customer and cost-per-order not yet formal engineering targets
  • Storage lifecycle tiering not enforced across all buckets; product telemetry tiering planned but not implemented
  • Infracost PR integration not enforced across all product repositories
  • Legacy manually provisioned resources persist in two older analytics accounts
  • CI runner spot fallback still manual
  • Spot expansion targets not yet approved
  • Low-severity anomaly escalation timing not defined in the runbook
  • Several development accounts lack team-level Slack routing for anomaly alerts
  • Idle resource waste at 4.8% of monthly spend against a below-3% 2026 target
  • Commitment coverage at 68% vs. 72% target for stable compute and database workloads
  • Exit-cost analysis for proprietary managed services performed only for tier-1 systems
  • Domain F (GenAI & AI Cost Management): 0/15 — no source evidence present

Confirmed anti-patterns

  • 6 confirmed anti-patterns present (specific anti-pattern labels not enumerated in Phase 2 output)
  • 19 verified anti-pattern absences confirmed
  • 5 anti-patterns not assessable from available source evidence

Verified anti-pattern absences

  • [A1] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The raw material shows the opposite of tag sprawl: mandatory tag keys, daily defect reporting, remediation SLA, exception expiry, and policy-as-code checks for new Terraform modules. No evidence supports significant untagged spend, ad-hoc taxonomy, or acceptance of untagged spend as overhead. Coverage interpretation: Relevant tagging governance and remediation coverage is present in src-001-c001 and src-004-c001, so this anti-pattern would likely have been revealed if present.
  • [A2] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source contradicts black-box cloud spend: engineering has resource-level dashboard views and drill-downs, dashboards refresh every 12 hours, and alerts route to owning teams. No evidence suggests cost visibility is restricted to finance or a single admin. Coverage interpretation: Dashboard and alert-routing evidence in src-002-c001 directly covers engineering visibility, making absence meaningful.
  • [A3] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source supports 12-hour dashboard refreshes and daily automated anomaly detection with Slack/on-call routing. There is no evidence of reporting delays over 48 hours, manual spreadsheet-based reporting, or lack of automated anomaly alerting. Coverage interpretation: Cost reporting cadence and anomaly automation are explicitly covered in src-002-c001, so the absence of delayed reporting is reasonably tested.
  • [A4] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source describes a unified FinOps dashboard integrating AWS, Azure, and GCP billing exports with role-appropriate views. There is no evidence of isolated provider/team views or finance and engineering using conflicting tools or numbers. The unit-economics gap does not by itself evidence siloed cost views. Coverage interpretation: Cross-cloud dashboard integration and shared stakeholder views are explicitly covered in src-002-c001, so absence of provider/account silos is meaningfully tested.
  • [B1] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The raw material contradicts commitment avoidance: it documents active commitment management, defined target coverage, current coverage, monthly utilization review, escalation of unused commitments, committed-discount utilization tracking, and quarterly strategy review. No source evidence supports fear-driven avoidance or predominantly on-demand use for stable workloads. Coverage interpretation: Commitment management is directly covered in src-003-c001 and src-002-c001. The evidence is specific enough that a commitment-avoidance pattern would likely be visible if present.
  • [B3] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source shows non-production shutdown outside business hours, owner-approved exceptions, weekly detection of idle/orphaned resources, and waste reporting with a reduction target. This contradicts the core anti-pattern conditions of dev/test running permanently and no lifecycle management or cleanup ownership. Coverage interpretation: Waste controls are directly covered in src-003-c001. Although residual waste exists, the source includes controls that would reasonably reveal hoarding/no-cleanup behavior if present.
  • [C1] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. No source chunk evidences personal-card spend, unlinked accounts, unknown cloud exposure, or cloud accounts outside the governed structure. The noted legacy manually provisioned resources and development accounts without Slack routing are governance gaps, not direct shadow IT evidence. Coverage interpretation: The packet includes relevant governance, tagging, billing-dashboard, account-dimension monitoring, and known account limitation material. These sources would reasonably surface unmanaged account exposure if documented in the available evidence.
  • [C2] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source material shows active budget governance: monthly budget-to-actual comparison, variance flagging above 10%, forecasting inputs, monthly reviews, and corrective actions. This contradicts budget blowout tolerance. Coverage interpretation: Budgeting and forecasting are directly covered in src-001-c001 and src-004-c001, with enough process detail to assess whether overruns are ignored or disconnected from operations.
  • [C3] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The source material contains measurable FinOps outcomes, including accepted right-sizing recommendations, EUR 46,000 monthly run-rate reduction, and an anomaly remediation with EUR 18,400 estimated monthly avoided cost. This contradicts the FinOps Theater anti-pattern. Coverage interpretation: The packet includes operating model, cadence, optimization review, and anomaly remediation evidence, which is relevant and sufficient to test whether FinOps activities produce measurable outcomes.
  • [C5] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source material shows data residency requirements reviewed during architecture design and compliance-driven costs such as logging, encryption, and regional replication tracked separately. No source evidence shows compliance cost surprises or compliance treated reactively. Coverage interpretation: Compliance cost handling is directly covered in src-004-c001, and architecture standards also include data-retention cost impact. This is enough to test the stated anti-pattern in the available evidence.
  • [D2] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The source directly contradicts cost-blind architecture: architecture review templates require estimated monthly cost, scaling cost curves, data-retention cost impact, and cost-performance tradeoff sections. Infracost is also visible in some pull requests, although not enforced everywhere. No harmful cost-blind decision pattern is evidenced. Coverage interpretation: Architecture governance and deployment-cost evidence are directly covered in src-004-c001, and the documented requirements are the opposite of the anti-pattern criteria.
  • [D4] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The evidence contradicts single-cloud tunnel vision: the dashboard integrates AWS, Azure, and GCP billing exports, commitments are managed across all three providers, and vendor pricing is benchmarked during annual renewal. The limited exit-cost analysis for proprietary services is a vendor-risk gap, not evidence of single-cloud provider loyalty. Coverage interpretation: The packet has direct multi-cloud cost-management and vendor-benchmarking evidence, which would reasonably reveal this anti-pattern if the organization were single-cloud or avoided provider comparison.
  • [E1] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The raw source describes the opposite of 'Cost is IT's Problem': product domains appoint engineering cost owners, product teams own cost decisions for their services, team leads receive showback, and anomaly alerts route to owning teams. No harmful-pattern criterion is evidenced. Coverage interpretation: Coverage is directly relevant to ownership and accountability. The operating model, showback policy, and anomaly routing would reasonably reveal if cost were isolated to IT; instead they show distributed accountability across Finance, FinOps, Platform Engineering, and product teams.
  • [E2] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source does not show blame, punishment, hidden cost problems, or adversarial finance-engineering interactions. The anomaly example shows routed alerting, correction by the owning team, follow-up architecture review, and recorded avoided cost. Corrective actions are also captured in a FinOps action log. Coverage interpretation: Coverage includes an anomaly runbook, a recent anomaly response, budget variance process, and cross-functional review cadence. These are relevant enough to test for blame-based behavior, and the described behavior is collaborative rather than punitive.
  • [E3] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source supports a staffed FinOps function, a federated cost-owner model, recurring reviews, dashboards/anomaly tooling, optimization backlog processes, and executive steering engagement. That contradicts the main FinOps Lip Service indicators of no dedicated team, no structured program, and optimization expected to happen organically. Explicit training budget is not shown, but no harmful criterion is positively evidenced. Coverage interpretation: Coverage includes operating model, headcount, tooling, governance cadence, and optimization processes. These materials would reasonably expose whether FinOps were merely nominal; instead they show operational implementation.
  • [E4] Tested absent: Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source shows shared cross-functional forums and tooling rather than a finance-engineering wall: the FinOps Council includes Finance, Platform Engineering, Security, and Product Operations; monthly reviews bring Finance and engineering leads together; and the dashboard includes executive, finance, and engineering views. No harmful-pattern criterion is evidenced. Coverage interpretation: Coverage is directly relevant to finance-engineering collaboration, tooling, and governance forums. Although individual financial/technical literacy is not assessed, the available source shows structures that bridge rather than separate finance and engineering.
  • [F3] Tested absent: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. There are no references to AI models, premium model usage, cheaper model alternatives, AI caching/batching, prompt reduction, retries, or agent loops. Premium Model Overuse is not evidenced. Coverage interpretation: The source is silent on AI model usage and model-selection practices, so the anti-pattern cannot be tested absent.
  • [F4] Tested absent: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The documents contain no RAG, prompt, memory, context-window, conversation-history, retrieval-scope, cache-efficiency, or token-budget evidence. Unbounded Context Growth is not evidenced. Coverage interpretation: The source does not cover AI context or retrieval architecture. Absence is therefore unknown rather than tested absent.
  • [F5] Tested absent: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. No AI adoption metrics, AI pilot reporting, AI activity dashboards, AI ROI claims, or continued AI use cases with unclear value are present in the source. AI Value Theater is not evidenced. Coverage interpretation: The documents are silent on AI adoption/value reporting, so the anti-pattern cannot be meaningfully tested absent.

Anti-patterns not assessable from source

  • [A5] Not assessed: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 1 to 0. Verifier status: unsupported. Adjudication: The cited sources show limited unit economics maturity and some savings reported as absolute EUR run-rate reduction, but they do not show misleading, unverifiable, presentation-only, or non-actionable vanity metrics. The optimization evidence indicates operational use through right-sizing recommendations, product-team approval/rejection, and backlog decisions. Coverage interpretation: Source coverage discusses unit economics gaps and optimization reporting, but does not provide enough targeted evidence about metric definitions, dashboard intent, or savings-verification practices to support a harmful vanity-metrics finding.
  • [B5] Not assessed: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: weak. The source supports commitment utilization tracking and multi-cloud commitment management, which argues against a total absence of rate optimization governance. However, it does not specifically evidence discount-interaction understanding, overlap analysis, or avoidance of double-coverage. The scanner's zero positive score is acceptable, but its claim that the source directly contradicts all discount-stacking ignorance criteria is overstated. Coverage interpretation: The available commitment evidence is relevant but not detailed enough to prove absence of discount-overlap or stacking-governance issues. Targeted evidence on overlap analysis and discount interaction governance would be needed.
  • [E5] Not assessed: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The source does not positively evidence a static maturity assumption. It shows ongoing optimization activity, weekly recommendations, monthly/quarterly review cadences, waste targets, and measured run-rate reduction. However, the source is also silent on formal FinOps maturity assessments and external maturity benchmarking, so absence of the anti-pattern is not fully testable. Coverage interpretation: The available material meaningfully contradicts stalled optimization momentum, but it does not fully cover maturity-assessment or external benchmarking practices. Therefore the harmful pattern is not evidenced, but full absence cannot be confirmed.
  • [F1] Not assessed: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The scanner correctly does not score Invisible Token Spend as present. The source does not establish that AI/API/token/model spend exists at all, so invisible AI spend cannot be confirmed from these documents. Coverage interpretation: The documents cover general cloud FinOps governance, but are silent on AI consumption. Because there is no evidence of existing AI spend, absence of the anti-pattern is not tested; it is unknown/not assessable.
  • [F2] Not assessed: Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. No AI experiments, playgrounds, sandbox AI keys, prototypes, or AI production transitions are described. The scanner correctly leaves the anti-pattern unscored rather than asserting absence. Coverage interpretation: The supplied material is silent on AI lifecycle or GenAI workload promotion. This does not prove the anti-pattern is absent; it only makes it not assessable from the provided evidence.

Silent / missing evidence

  • Domain F (GenAI & AI Cost Management): 0/15 — complete evidence gap; no AI or token-cost governance documented
  • 7 silent areas across domains where no source evidence was available
  • Actual tag compliance rates at resource level not evidenced (policy and reporting process documented only)
  • Adoption rate of architecture review templates across all services not evidenced
  • Operational observance of documented weekly and monthly cadences not independently verified
  • Scope and spend contribution of two legacy analytics accounts not quantified in source

Evidence summary for the FinOps Lead

Current-State Snapshot: The assessed organization scores 50/100 on the evidence-gated FinOps Readiness scale, classified as Walk maturity. The Maturity Depth Index is 47%, indicating that while walk-level practices are present, they are not uniformly deep or enforced. Anti-Pattern Burden is confirmed at 7% across 6 confirmed anti-patterns, with Anti-Pattern Clearance at 63% and Coverage at 83%. Seven silent areas and seven maturity gaps remain unresolved. Confirmed gaps include: chargeback not yet active (showback only); unit economics formal targets limited to checkout transactions and loyalty API calls; storage lifecycle tiering not enforced across all buckets; Infracost PR integration not enforced across all product repositories; CI runner spot fallback still manual; spot expansion targets not yet approved; idle resource waste at 4.8% of monthly spend against a below-3% target; commitment coverage at 68% against a 72% target; and low-severity anomaly escalation timing not yet defined in the runbook. Domain F (GenAI & AI Cost Management) scored 0/15 with no source evidence.

Source Confidence & Boundaries: The source documents are policy and operational documents, not audit telemetry or system exports. The audit can confirm what is documented as policy or process intent; it cannot independently verify enforcement rates, actual tagging compliance percentages, or whether documented cadences are operationally observed. Gaps between documented policy and confirmed enforcement represent the primary evidence boundary.

Confidence Notes — Unverified Claims

The following statements could not be verified against the source after 3 regenerate pass(es). Treat with caution.

Evidence summary for the CFO

Current-State Snapshot: The assessed organization is classified at Walk maturity with an evidence-gated readiness score of 50/100 and a Maturity Depth Index of 47%. Seven silent areas remain where no evidence was available.

Evidence-Backed Findings: The governance framework includes a FinOps Council with Finance, Platform Engineering, Security, and Product Operations representation, monthly budget variance reviews with a 10% variance flag threshold, and reporting to the CTO/CFO steering group. Each product team holds a monthly cloud budget with actuals comparison. A documented corrective-action log captures budget responses. Financially material gaps include: chargeback not yet active, limiting cost accountability signals to showback only; unit economics defined only for two transaction types (checkout transactions and loyalty API calls), with cost-per-customer and cost-per-order not yet formal engineering targets; idle resource waste confirmed at 4.8% of monthly spend against a 2026 target of below 3%; commitment coverage at 68% against a 72% target for stable compute and database workloads; and storage lifecycle tiering not enforced across all buckets. A single confirmed anomaly event (checkout-api compute, 2026-03-18) documented EUR 18,400 estimated monthly avoided cost after correction. Quarterly commitment and architecture cost reviews with the CTO/CFO steering group are documented. Domain F (GenAI & AI Cost Management) returned a score of 0/15, representing a complete evidence gap for AI-related financial controls.

Source Confidence & Boundaries: Financial figures cited (EUR 46,000 monthly run-rate reduction from right-sizing; EUR 18,400 anomaly avoided cost) are sourced from internal operational documents, not independently audited financial statements. The audit cannot confirm whether showback views are acted upon by product teams or whether budget variance responses result in durable spend changes. Chargeback pilot plans (2026-Q3) are documented as plans, not confirmed outcomes.

Confidence Notes — Unverified Claims

The following statements could not be verified against the source after 3 regenerate pass(es). Treat with caution.

Evidence summary for the Engineering Lead

Current-State Snapshot: The assessed organization sits at Walk maturity (50/100, Maturity Depth Index 47%). Anti-Pattern Burden is confirmed at 7% with 6 confirmed anti-patterns.

Evidence-Backed Findings: Confirmed engineering strengths include: Terraform as the standard for production cloud infrastructure with mandatory tags and policy-as-code checks in new modules; Infracost estimates visible in pull requests for the platform account; architecture review templates requiring estimated monthly cost, scaling cost curve, data-retention cost impact, and a cost-performance tradeoff section; autoscaling policies required to define maximum capacity, target utilization, and cost ceiling guardrails; engineering teams able to view service cost before and after releases; non-production environment shutdown policies outside business hours unless tagged always_on=true; and weekly waste detection for idle load balancers, unattached disks, and orphaned snapshots. Confirmed engineering gaps include: Infracost PR integration not yet enforced across all product repositories; legacy manually provisioned resources persisting in two older analytics accounts; CI runner spot fallback still manual (analytics batch interruption handling documented as present); storage lifecycle tiering for product telemetry planned but not enforced; idle resource waste at 4.8% of monthly spend against a below-3% target; and commitment coverage at 68% against a 72% target. Spot instances are approved for batch analytics, CI test runners, and image-processing jobs, but expansion targets are not yet approved and CI runner fallback remains manual. Exit-cost analysis for proprietary managed services is performed only for tier-1 systems.

Source Confidence & Boundaries: Source documents describe policies and engineering standards as written. The audit can confirm what standards are documented; it cannot confirm adoption rates in product repositories, actual tag compliance rates at resource level, or whether architecture review templates are consistently applied to all new services. The two analytics accounts with legacy manual provisioning are identified by type but not quantified in scope or spend contribution.

Diagnosis

Interpretation of evidence — not the implementation plan

Root causes

  • Manually provisioned legacy resources in two analytics accounts fall outside Terraform policy-as-code controls — directly evidenced in source
  • Showback-only allocation is documented; the mechanism by which this affects product team financial accountability behavior is not evidenced (evidence gap)
  • Spot expansion scope is constrained: targets not yet approved, and CI runner fallback remains manual — directly evidenced in source
  • Storage lifecycle tiering enforcement gap for product telemetry is documented as planned but not implemented — directly evidenced in source
  • Low-severity anomaly escalation timing undefined in the runbook — directly evidenced in source
  • Domain F evidence absence: whether the organization has GenAI workloads requiring cost governance is not stated in the source; the 0/15 score reflects a documentation gap, not a confirmed operational failure

Domain diagnosis

  • A: Cost Visibility & Allocation: Score 9/15. Multi-cloud dashboard with 12-hour refresh, tag-based drill-down, and showback to product teams are confirmed strengths. Tagging policy is documented with daily defect reporting and remediation SLA. Gaps include showback-only allocation (chargeback not yet active) and unit economics coverage limited to two transaction types. Actual tag compliance rates are not evidenced beyond the policy and reporting mechanism.
  • B: Rate & Usage Optimization: Score 8/15. Commitment management across Reserved Instances, Savings Plans, and committed-use discounts is documented with a 72% coverage target; current coverage is 68%. Monthly utilization review and EUR 5,000 escalation threshold for unused commitments are documented. Right-sizing recommendations accepted in 2026-Q1 reduced monthly run-rate by EUR 46,000. Spot is approved for batch analytics, CI runners, and image-processing, but expansion targets are not yet approved and CI runner fallback is manual. Idle waste at 4.8% exceeds the below-3% target.
  • C: Governance & Policy: Score 9/15. FinOps Council with cross-functional membership, monthly budget variance reviews with a 10% flag threshold, and a documented corrective-action log are confirmed. Tagging exception governance with 30-day expiry is documented. Gaps include undefined low-severity anomaly escalation timing and the absence of chargeback governance. Exit-cost analysis for proprietary managed services is scoped only to tier-1 systems.
  • D: Architecture & Engineering: Score 8/15. Architecture review templates with cost and scaling requirements, autoscaling guardrails, Terraform-as-standard with policy-as-code, and Infracost in platform-account PRs are confirmed. Confirmed gaps: Infracost not enforced across all product repositories; legacy manually provisioned resources in two analytics accounts; storage lifecycle tiering for product telemetry not enforced; CI runner spot fallback manual.
  • E: Culture & Organization: Score 8/15. Federated model with engineering cost owners per product domain, weekly optimization backlog reviews, and a documented RACI are confirmed. Engineering teams have access to pre- and post-release cost views. Gaps include development accounts without team-level anomaly Slack routing and unit economics not yet used as formal engineering targets beyond two transaction types. Whether the operating model cadences are consistently observed is not independently verifiable from source documents.
  • F: GenAI & AI Cost Management: Score 0/15. No source evidence addresses GenAI workloads, AI/ML token cost governance, model inference cost allocation, or related controls. The audit cannot determine whether this reflects an absence of GenAI workloads or an absence of documentation; both remain evidence gaps.

Confidence (medium): Evidence Density is 57%, below the threshold that would lift the readiness cap. Source documents are policy and operational documents rather than audit telemetry or system exports, meaning documented practices cannot be independently verified as operationally enforced. Quantitative metrics (waste percentage, commitment coverage, run-rate reduction) are taken from internal documents at face value. Domain F is entirely unassessed. These factors collectively limit diagnosis confidence to medium.

Planning Decision: CONDITIONAL GO

Locked findings support action on several confirmed gaps (Infracost PR coverage, storage lifecycle enforcement, waste reduction, commitment coverage, low-severity anomaly escalation, chargeback progression, unit economics extension), but evidence density is 57% with 7 silent areas and Domain F entirely unassessed. Proceed with grounded remediation while validating silent areas and confirming whether GenAI workloads exist before prescribing AI-specific tactics.

Safe to act on

  • Extending Infracost PR integration beyond the platform account to product repositories (confirmed enforcement gap)
  • Enforcing storage lifecycle tiering for product telemetry (confirmed planned-but-not-implemented gap)
  • Progressing waste reduction toward the below-3% target from the confirmed 4.8% baseline
  • Progressing commitment coverage toward the 72% target from the confirmed 68% baseline
  • Defining low-severity anomaly escalation timing in the runbook (confirmed runbook gap)
  • Extending team-level Slack routing for anomaly alerts to development accounts currently lacking it
  • Preparing for the documented 2026-Q3 chargeback pilot for shared analytics workloads
  • Bringing legacy manually provisioned resources in the two older analytics accounts under Terraform policy-as-code

Evidence needed before action

  • Confirmation of whether GenAI workloads exist in the assessed organization before applying Domain F tactics
  • Actual tag compliance rates at resource level to validate the documented policy
  • Adoption rate of architecture review templates across product services
  • Independent verification of documented weekly and monthly cadence observance
  • Quantified scope and spend contribution of the two legacy analytics accounts
  • Evidence on the five not-assessable anti-patterns and seven silent areas

Source Coverage Gaps

To strengthen the next assessment cycle, include the following kinds of evidence in the source document.

other

Remediation Roadmap

1. Crawl — Foundation (0-3 Months)

Why

The locked findings show a Walk-level program with a 57% evidence density that caps readiness at 50/100 and leaves seven silent areas including all of Domain F. Before scaling controls, the assessed organization must close visible runbook and routing gaps that are directly evidenced: undefined low-severity anomaly escalation timing, development accounts without team-level Slack routing, and the documentation absence for GenAI cost management. These are low-risk foundational fixes that improve evidence density and unlock later phases. Validation activity for silent areas is prioritized so subsequent optimization is grounded in verified current state rather than assumed enforcement.

What

The intended change is to close the confirmed near-term visibility and routing gaps and to establish the evidence base needed for later phases. This means completing the anomaly runbook, extending team-level alert routing where it is missing, and confirming whether GenAI workloads exist so Domain F can either be scoped or set aside. The organization should also validate documented practices with operational evidence to raise evidence density above the cap threshold.

How

2. Walk — Optimization (3-6 Months)

Why

Locked findings evidence multiple enforcement gaps that are ready for action: Infracost PR integration is limited to the platform account, storage lifecycle tiering for product telemetry is planned but not implemented, idle resource waste sits at 4.8% against a below-3% target, and commitment coverage is 68% against a 72% target.

What

The intended change is measurable movement toward existing documented targets rather than the creation of new baselines. Infracost PR checks should extend beyond the platform account into product repositories; storage lifecycle policy should be operationalized for product telemetry; waste reduction should progress toward the below-3% target; and commitment coverage should progress toward the 72% target. Each action is scoped to the exact gap identified in the locked findings and uses the existing governance forums (FinOps Council, weekly optimization backlog review, quarterly commitment review) as the accountability mechanism.

How

3. Walk — Embedding (6-12 Months)

Why

The locked findings identify allocation and accountability limits that require more time to embed: chargeback is not yet active with a pilot documented for 2026-Q3, unit economics are formalized only for checkout transactions and loyalty API calls, and CI runner spot fallback remains manual with spot expansion targets not yet approved.

What

The intended change is to progress the documented 2026-Q3 chargeback pilot for shared analytics workloads and to extend unit economics beyond the two currently formalized transaction types where the FinOps Council determines it is appropriate. Manual CI runner spot fallback should move to an automated pattern within the currently approved spot scope. Spot expansion targets should be brought through the existing quarterly review for approval. The phase ends when allocation and optimization mechanisms are consistently operating rather than partially implemented.

How

4. Run — Continuous (12+ Months)

Why

The locked findings note that exit-cost analysis for proprietary managed services is performed only for tier-1 systems and that Domain F returned 0/15 with no source evidence. If phase-1 validation confirms GenAI workloads exist, AI cost governance becomes a grounded Run-phase priority; if not, the domain remains a documentation gap rather than an operational failure. The Run phase also sustains the outcome-tracking discipline needed to keep the improvements from phases 2 and 3 from decaying, using the existing FinOps Council and steering-group cadence as the accountability structure.

What

The intended change is durable operation of the mechanisms established earlier, with two additions grounded in locked findings: extending exit-cost analysis beyond tier-1 systems where the FinOps Council prioritizes it, and either establishing AI cost governance (conditional on phase-1 GenAI validation) or formally recording Domain F as not applicable. Outcome tracking should keep the FinOps program measured against documented targets rather than activity volume. No new current-state assertions are introduced; the phase depends on prior evidence collection and successful phase-2 and phase-3 execution.

How

Forensic Audit: FinOps Maturity

A · Cost Visibility & Allocation

A1

Comprehensive Cost Allocation & Tagging

OK

The organization must implement a consistent, enforced tagging and labeling strategy across all cloud resources. Tags must map spend to business units, applications, environments, and cost centers. Untagged resources must be actively tracked and remediated.

AI Reasoning

Crit 1: Found. A documented and enforced tagging policy mapping resources to business_unit, product, environment, owner_team, cost_center is explicitly described with enforcement via policy-as-code in Terraform. Crit 2: Found. Untagged/mis-tagged resources are reported daily in the FinOps dashboard and Platform Engineering must remediate within five business days (defined SLA). Crit 3: Found. The tagging strategy explicitly supports the showback model (weekly cost allocation to product teams by tags) and the optimization backlog review cadence. All 3 met. Total: 3.

Evidence
  • “All production cloud resources must carry business_unit, product, environment, owner_team, cost_center, and data_classification tags. Untagged or mis-tagged resources are reported daily in the FinOps dashboard. Platform Engineering must remediate critical tag defects within five business days. Exceptions require FinOps Council approval and expire after 30 days.” — Tagging policy · Policy
  • “Cloud costs are allocated weekly to product teams using tags, account hierarchy, and shared platform allocation rules. Team leads receive showback views for product, environment, and owner_team.” — Showback model · Accountability
  • “New Terraform modules include mandatory tags and policy-as-code checks. Infracost estimates are visible in pull requests for the platform account, but not yet enforced for all product repositories.” — Infrastructure as Code · Automation
A2

Showback & Chargeback Models

Partial

The organization must have mechanisms to attribute cloud costs back to consuming teams or business units. This can be showback (visibility without billing) or full chargeback (actual cost transfer). The model must be transparent and accepted by stakeholders.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 1 to 1 after a targeted rescan. Verifier status: supported. The source supports a functioning weekly showback model using tags, account hierarchy, and shared platform allocation rules. The scanner correctly withheld credit for formal stakeholder acceptance and regular refinement of the allocation model, which are not evidenced.

Evidence
  • “Cloud costs are allocated weekly to product teams using tags, account hierarchy, and shared platform allocation rules. Team leads receive showback views for product, environment, and owner_team.” — Showback model · Accountability
  • “Chargeback is not yet active; Finance plans to pilot chargeback for shared analytics workloads in 2026-Q3.” — Showback model · Policy
A3

Cost Anomaly Detection & Alerting

Partial

The organization must have automated systems to detect unexpected cost spikes, billing anomalies, and budget threshold breaches. Alerts must reach the right stakeholders in near-real-time with actionable context.

AI Reasoning

Crit 1: Found. Automated anomaly detection runs daily, well within the 24-hour threshold. Crit 2: Found. Alerts include affected service, account, owner_team, amount, percentage variance, and route to the owning team Slack channel — actionable context is present. Crit 3: Partial/Not fully met. Escalation is defined only for high-severity anomalies (one business day acknowledgement); the runbook explicitly lacks escalation timing for low-severity anomalies, and development accounts lack team-level routing. This prevents a score of 3. Total: 2.

Evidence
  • “Automated anomaly detection runs daily against account, product, service, and environment dimensions. Alerts include affected service, account, owner_team, amount, percentage variance, and a link to the dashboard. Alerts route to the owning team Slack channel and to the FinOps on-call analyst. High-severity anomalies require acknowledgement within one business day.” — Anomaly detection · Automation
  • “On 2026-03-18, checkout-api compute spend increased 31% after a misconfigured autoscaling threshold. The alert was routed to Platform Checkout and FinOps. The team corrected the scaling policy, created a follow-up architecture review, and recorded estimated monthly avoided cost.” — Recent anomaly example · Process
  • “The anomaly runbook does not yet define escalation timing for low-severity anomalies. Several development accounts are monitored daily but do not yet have team-level Slack routing.” — Known limitation · Process
A4

Cloud Cost Dashboards & Reporting

Partial

The organization must provide accessible, role-appropriate cost dashboards to all stakeholders. Engineering sees resource-level detail; finance sees business-unit rollups; executives see trends and forecasts. Dashboards must be current, not stale.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 2 after a targeted rescan. Verifier status: supported. The dashboard evidence supports role-appropriate executive, finance, and engineering views, 12-hour refresh, trend/forecast information, and AWS/Azure/GCP billing export integration. The scanner correctly withheld full credit because SaaS platform integration is not evidenced.

Evidence
  • “The FinOps dashboard refreshes every 12 hours and integrates AWS, Azure, and GCP billing exports. It contains executive trend views, finance budget views, and engineering resource-level views.” — Dashboards · Operational
  • “The dashboard tracks month-to-date spend, forecasted month-end spend, previous-period comparison, and committed-discount utilization.” — Dashboards · Operational
A5

Unit Economics & Cost-Per-Transaction

Partial

The organization must measure cloud cost efficiency in business terms — cost per customer, per transaction, per API call, or per revenue dollar. Raw spend is insufficient; unit economics connect cloud investment to business value.

AI Reasoning

Crit 1: Partial. Unit economics exist for only two metrics (checkout transactions and loyalty API calls); coverage is incomplete and explicitly acknowledged as a gap. Score capped at 1 for partial/aspirational state. Crit 2: Not found. No evidence of unit cost trends being monitored over time to separate growth-driven spend from efficiency degradation. Crit 3: Not found. Engineering teams are explicitly NOT given unit cost targets as formal metrics. Total: 1.

Evidence
  • “Unit economics are defined only for checkout transactions and loyalty API calls. Cost per customer and cost per order are not yet used as formal engineering targets.” — Policy gaps · Policy

B · Rate & Usage Optimization

B1

Commitment-Based Discounts (RIs, Savings Plans, CUDs)

Partial

The organization must actively manage commitment-based discount instruments (Reserved Instances, Savings Plans, Committed Use Discounts) to reduce on-demand pricing. Coverage targets must be set, tracked, and optimized.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 2 after a targeted rescan. Verifier status: supported. The cited chunks support active RI/Savings Plan/CUD management across AWS/Azure/GCP, a defined 72% coverage target, current 68% coverage tracking, monthly utilization review, escalation of unused commitments, dashboard utilization tracking, and quarterly commitment strategy review. The source does not show exchange/resale/remediation of unused commitments, so a score of 2 is appropriate.

Evidence
  • “Northstar Retail manages Reserved Instances, Savings Plans, and committed-use discounts across AWS, Azure, and GCP. Target commitment coverage is 72% for stable compute and database workloads. Current coverage is 68%.” — Commitment management · Operational
  • “Utilization is reviewed monthly; unused commitments above EUR 5,000 monthly amortized value are escalated to Finance and Platform Engineering.” — Commitment management · Process
  • “Quarterly: commitment strategy and architecture cost review with CTO/CFO steering group.” — Cadence · Process
  • “The dashboard tracks month-to-date spend, forecasted month-end spend, previous-period comparison, and committed-discount utilization.” — Dashboards · Operational
B2

Right-Sizing & Workload Optimization

Partial

The organization must continuously analyze resource utilization and right-size instances, databases, and services to match actual demand. Over-provisioned resources represent waste; under-provisioned resources create performance risk.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 2 after a targeted rescan. Verifier status: supported. The source supports right-sizing recommendations generated weekly from CPU, memory, and storage utilization, product-team review in an optimization backlog, weekly backlog review, and tracked financial impact of EUR 46,000 monthly run-rate reduction. Automatic generation is not explicit, so criterion 2 is only partially supported; count 2 is reasonable.

Evidence
  • “Recommendations are generated weekly from CPU, memory, and storage utilization. Product teams must approve or reject recommendations in the cost optimization backlog.” — Optimization actions · Process
  • “During 2026-Q1, engineering accepted 22 right-sizing recommendations and deferred 7 due to release freeze risk. Accepted actions reduced monthly run-rate by EUR 46,000.” — Optimization actions · Financial-Integration
  • “Weekly: optimization backlog review with product cost owners.” — Cadence · Process
B3

Waste Elimination (Orphaned Resources & Idle Instances)

Partial

The organization must actively identify and terminate orphaned resources (unattached volumes, unused IPs, idle load balancers), zombie instances, and dev/test environments left running outside business hours.

AI Reasoning

Crit 1: Found. Idle load balancers, unattached disks, and orphaned snapshots detected weekly and acted upon — systematic identification and termination is described. Crit 2: Found. Non-production environments shut down outside business hours with an automated tag-based exception mechanism (always_on=true with owner approval). Crit 3: Partially found. Waste is measured as a percentage of monthly cloud spend (4.8%) with a stated 2026 target (below 3%), satisfying reporting and target aspects. However, this is a single aggregate metric and does not provide explicit evidence of a broader multi-dimensional waste reporting system — but the criteria are substantially met. Awarding Crit 3 as met given the waste percentage and reduction target are explicitly stated. Total: 3.

Evidence
  • “Non-production environments shut down outside business hours unless tagged always_on=true with owner approval. Idle load balancers, unattached disks, and orphaned snapshots are detected weekly.” — Waste controls · Automation
  • “The waste report shows idle resource waste at 4.8% of monthly cloud spend, against a 2026 target of below 3%.” — Waste controls · Accountability
B4

Spot & Preemptible Instance Strategy

Partial

The organization must leverage spot/preemptible instances for fault-tolerant and batch workloads to achieve significant rate reductions. Workload architecture must support interruption handling.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 1 to 1 after a targeted rescan. Verifier status: supported. The source supports spot/preemptible approval for batch analytics, CI test runners, and image-processing jobs. It also explicitly limits maturity: interruption handling exists only for analytics batch jobs, CI fallback remains manual, and expansion targets are not approved. Count 1 is supported.

Evidence
  • “Spot instances are approved for batch analytics, CI test runners, and image-processing jobs.” — Spot and preemptible usage · Operational
  • “Interruption handling exists for analytics batch jobs, but CI runner fallback is still manual. Spot savings are tracked, but expansion targets are not yet approved.” — Spot and preemptible usage · Operational
B5

Storage Lifecycle & Data Tiering

Partial

The organization must implement storage lifecycle policies that automatically tier data from hot to warm to cold storage based on access patterns. Retention policies must be enforced to prevent unbounded storage growth.

AI Reasoning

Crit 1: Partially found. Archive tiering exists for audit logs older than 180 days — partial hot/warm/cold tiering — but product telemetry lifecycle tiering is 'planned but not yet enforced' and tiering is not enforced across all buckets. At best this is one partial implementation. Crit 2: Partially found. Retention is defined for security, audit, and application logs, but enforcement is incomplete (not across all buckets). Crit 3: Not found. No evidence of monitoring storage cost per GB or optimizing for appropriate storage class. Given plans and partial implementation but no enforcement breadth and no storage-cost-per-GB monitoring, total: 1 (aspirational/partial).

Evidence
  • “Log retention is defined for security, audit, and application logs. Archive tiering exists for audit logs older than 180 days. Object storage lifecycle tiering for product telemetry is planned but not yet enforced.” — Storage · Policy
  • “Storage retention rules exist for logs, but lifecycle tiering is not yet enforced across all buckets.” — Policy gaps · Policy

C · Governance & Policy

C1

Cloud Financial Policy Framework

Partial

The organization must have a documented cloud financial management policy framework covering spend limits, approval workflows, resource provisioning standards, and accountability structures. Policies must be enforced, not aspirational.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 2 after a targeted rescan. Verifier status: supported. The source supports a documented governance framework, tagging requirements, exception approval, named governance ownership, and some policy-as-code enforcement in Terraform modules. It also supports the scanner's caveats: spend limits/full provisioning approval workflows and policy lifecycle review cadence are not explicit.

Evidence
  • “The FinOps Council owns cloud financial governance with named representatives from Finance, Platform Engineering, Security, and Product Operations.” — Governance · Policy
  • “All production cloud resources must carry business_unit, product, environment, owner_team, cost_center, and data_classification tags... Exceptions require FinOps Council approval and expire after 30 days.” — Tagging policy · Policy
  • “New Terraform modules include mandatory tags and policy-as-code checks.” — Infrastructure as Code · Automation
C2

Budget Management & Forecasting

Partial

The organization must maintain cloud budgets at the team, project, and organizational level with accurate forecasting. Budget-to-actual variance must be tracked monthly, with variance analysis driving corrective action.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 2 after a targeted rescan. Verifier status: supported. src-001-c001 directly supports monthly product-team budgets, monthly budget-to-actual variance checks above 10%, a forecasting model using historical spend/planned launches/commitment expiries/traffic growth, and corrective actions in a monthly FinOps action log. The scanner appropriately did not award full credit because project/org-level budget hierarchy is not explicit.

Evidence
  • “Each product team has a monthly cloud budget. Finance compares budget to actuals every month and flags variance above 10%.” — Budget and forecast · Process
  • “The forecasting model uses historical spend, planned product launches, reserved commitment expiries, and expected traffic growth.” — Budget and forecast · Process
  • “Corrective actions are captured in the monthly FinOps action log.” — Budget and forecast · Accountability
C3

FinOps Operating Model (Team, Roles, RACI)

OK

The organization must have a defined FinOps operating model with clear roles, responsibilities, and RACI matrices. Whether centralized, federated, or hybrid, the model must have executive sponsorship and operational cadence.

AI Reasoning

Crit 1: Found. A defined federated FinOps operating model with explicit named roles (FinOps lead, cloud economist, platform analyst, engineering cost owners) and a RACI matrix is documented. Crit 2: Found. Executive sponsorship is evidenced by the quarterly CTO/CFO steering group engagement and the FinOps Council reporting unresolved cost risks to CTO and CFO. Crit 3: Found. A documented operational cadence covering daily, weekly, monthly, and quarterly FinOps activities is explicitly stated. Total: 3.

Evidence
  • “The FinOps operating model is federated. The central FinOps team has one FinOps lead, one cloud economist, and one platform analyst. Each product domain appoints an engineering cost owner. The RACI states that Finance owns budgets, FinOps owns allocation and optimization process, Platform Engineering owns guardrails, and product teams own cost decisions for their services.” — Operating model · Accountability
  • “Quarterly: commitment strategy and architecture cost review with CTO/CFO steering group.” — Cadence · Process
  • “Daily: anomaly triage for high-severity alerts. Weekly: optimization backlog review with product cost owners. Monthly: budget variance review with Finance and engineering leads.” — Cadence · Process
C4

Procurement & Vendor Management

Partial

The organization must have a structured approach to cloud procurement including Enterprise Discount Programs (EDPs), negotiated pricing, and multi-year agreements. Vendor relationships must be actively managed, not passively consumed.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 1 to 1 after a targeted rescan. Verifier status: supported. The cited quote in src-004-c001 directly supports annual vendor pricing benchmarking. The scanner correctly treats RI/SP/CUD management as insufficient evidence for EDP/private pricing/negotiated rates and notes that exit-cost analysis is only partial.

Evidence
  • “Vendor pricing is benchmarked during annual renewal, but exit-cost analysis for proprietary managed services is only performed for tier-1 systems.” — Compliance and vendor management · Process
C5

Compliance & Regulatory Cost Controls

Partial

The organization must ensure cloud financial operations comply with regulatory requirements including data residency cost implications, audit trail requirements, and financial reporting standards.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 1 to 1 after a targeted rescan. Verifier status: supported. src-004-c001 supports data residency review during architecture design and separate tracking of compliance-driven cost increases. The scanner correctly withholds additional credit because audit trails for financial decisions and separate justification of compliance costs are not evidenced.

Evidence
  • “Data residency requirements are reviewed during architecture design. Compliance-driven cost increases such as dedicated logging, encryption, and regional replication are tracked separately in the Finance view.” — Compliance and vendor management · Process

D · Architecture & Engineering

D1

Cost-Aware Architecture Design

OK

Engineering teams must consider cost as a first-class architecture constraint alongside performance, reliability, and security. Architecture reviews must include cost modeling. The cheapest architecture that meets requirements wins.

AI Reasoning

Crit 1: Found. Architecture review templates include a cost-performance tradeoff section, placing cost as a first-class constraint alongside other factors. Crit 2: Found. Templates require estimated monthly cost and scaling cost curves — explicit cost modeling in architecture proposals. Crit 3: Found. Engineering teams can view service cost before and after releases, confirming cost data access during development. Total: 3.

Evidence
  • “Architecture review templates require estimated monthly cost, scaling cost curve, data-retention cost impact, and a cost-performance tradeoff section.” — Architecture standards · Process
  • “Engineering teams can view service cost before and after releases.” — Architecture standards · Operational
  • “Quarterly: commitment strategy and architecture cost review with CTO/CFO steering group.” — Cadence · Process
D2

Infrastructure as Code with Cost Guardrails

Partial

Infrastructure provisioning must be codified (Terraform, Pulumi, CloudFormation) with embedded cost guardrails. Policy-as-code must prevent provisioning of oversized or non-compliant resources before deployment.

AI Reasoning

Crit 1: Found. Terraform is the declared standard for production infrastructure provisioning. However, legacy manually provisioned resources exist in two accounts, so IaC is not universal — partial but clearly operational. Crit 2: Partially found. Policy-as-code checks exist in new Terraform modules, but enforcement is not complete across all repositories. Score is operational (2) not embedded (3). Crit 3: Found. Infracost estimates are visible in pull requests for the platform account, providing a cost estimation step in the CI/CD pipeline, though not yet enforced everywhere. Total: 2 (all three criteria partially met, but none fully enforced across the entire estate).

Evidence
  • “Terraform is the standard for production cloud infrastructure. New Terraform modules include mandatory tags and policy-as-code checks.” — Infrastructure as Code · Policy
  • “Infracost estimates are visible in pull requests for the platform account, but not yet enforced for all product repositories.” — Infrastructure as Code · Automation
  • “Legacy manually provisioned resources still exist in two older analytics accounts.” — Infrastructure as Code · Operational
D3

Automated Scaling Policies

Partial

The organization must implement automated scaling (horizontal and vertical) based on demand signals to avoid both over-provisioning (waste) and under-provisioning (performance degradation). Scaling policies must be tested and monitored.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 1 to 1 after a targeted rescan. Verifier status: supported. The cited evidence supports only limited autoscaling governance: autoscaling policies must define maximum capacity, target utilization, and cost ceiling guardrails, and one autoscaling misconfiguration was detected as a cost anomaly. There is no source evidence of systematic horizontal/vertical autoscaling implementation across production, scale-to-zero testing, or comprehensive scaling-event cost monitoring.

Evidence
  • “Autoscaling policies must define maximum capacity, target utilization, and cost ceiling guardrails.” — Architecture standards · Policy
  • “On 2026-03-18, checkout-api compute spend increased 31% after a misconfigured autoscaling threshold.” — Recent anomaly example · Operational
D4

Multi-Cloud & Hybrid Cost Management

Partial

If operating across multiple cloud providers or hybrid environments, the organization must have unified cost visibility, consistent tagging, and cross-provider optimization strategies. Single-pane-of-glass for total cloud economics.

AI Reasoning

Crit 1: Found. The FinOps dashboard integrates AWS, Azure, and GCP billing exports into a single unified view — this is a single-pane-of-glass for multi-cloud cost visibility. Crit 2: Partial. Tagging policy applies to production resources (src-001) and commitment management spans all three clouds (src-003), suggesting some consistency, but explicit cross-provider apples-to-apples tagging comparison is not described. Crit 3: Not found. No evidence that workload placement decisions are informed by cross-provider cost analysis. Vendor pricing is benchmarked at annual renewal (src-004) but this is not described as informing workload placement. Total: 2.

Evidence
  • “The FinOps dashboard refreshes every 12 hours and integrates AWS, Azure, and GCP billing exports. It contains executive trend views, finance budget views, and engineering resource-level views.” — Dashboards · Operational
  • “Northstar Retail manages Reserved Instances, Savings Plans, and committed-use discounts across AWS, Azure, and GCP.” — Commitment management · Operational
D5

Containerization & Serverless Optimization

NOK

The organization must optimize container orchestration (Kubernetes) for cost efficiency including pod right-sizing, cluster autoscaling, and bin-packing. Serverless workloads must be monitored for execution cost and duration optimization.

AI Reasoning

Final maturity assessment: NOK. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. No cited source mentions Kubernetes, pod right-sizing, node autoscaling, bin-packing, serverless function cost monitoring, execution-duration optimization, memory/concurrency optimization, or a decision framework for containers vs. serverless vs. VMs. The scanner’s zero score is supported by the available packet evidence.

E · Culture & Organization

E1

FinOps Team / Center of Excellence

OK

The organization must have a dedicated FinOps function — whether a central team, federated practitioners, or a virtual CoE. This team drives standards, tooling, training, and optimization campaigns across the organization.

AI Reasoning

Crit 1: Found. A dedicated federated FinOps team with named headcount, a charter (RACI), and a FinOps Council with cross-functional named representatives is explicitly described. Crit 2: Found. The team drives weekly optimization backlog reviews with product cost owners and runs quarterly commitment/architecture reviews; the Q1 review documents 22 accepted right-sizing actions and measurable run-rate reduction. Crit 3: Found. The operating model positions FinOps as owning 'allocation and optimization process' with engineering accountability embedded at the product-domain level, and engineering teams have pre/post-release cost visibility — indicating value-enabler positioning with engineering credibility. Total: 3.

Evidence
  • “The FinOps operating model is federated. The central FinOps team has one FinOps lead, one cloud economist, and one platform analyst. Each product domain appoints an engineering cost owner. The RACI states that Finance owns budgets, FinOps owns allocation and optimization process, Platform Engineering owns guardrails, and product teams own cost decisions for their services.” — Operating model · Accountability
  • “Weekly: optimization backlog review with product cost owners. Monthly: budget variance review with Finance and engineering leads. Quarterly: commitment strategy and architecture cost review with CTO/CFO steering group.” — Cadence · Process
  • “The FinOps Council owns cloud financial governance with named representatives from Finance, Platform Engineering, Security, and Product Operations. The council meets monthly, reviews cloud budget variance, and reports unresolved cost risks to the CTO and CFO steering group.” — Cloud Financial Governance Policy · Operational
E2

Engineering Cost Accountability

Partial

Engineering teams must own the cost of the resources they provision. Cost accountability must be embedded in team objectives, sprint reviews, and performance evaluations — not siloed in finance.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 1 after a targeted rescan. Verifier status: weak. The source clearly supports engineering visibility and accountability through service-cost visibility, engineering cost owners, and product-team approval/rejection of optimization recommendations. However, the second criterion is not supported because there is no evidence of cost efficiency in OKRs, sprint reviews, performance evaluations, or similar team objectives. The third criterion is only partially supported: engineers appear empowered to optimize, but the source does not show explicit incentives. The scanner's count of 2 is therefore too strong.

Evidence
  • “Engineering teams can view service cost before and after releases.” — Architecture standards · Operational
  • “Each product domain appoints an engineering cost owner.” — Operating model · Accountability
  • “Product teams must approve or reject recommendations in the cost optimization backlog.” — Optimization actions · Accountability
E3

Executive FinOps Sponsorship

Partial

FinOps must have executive sponsorship (CTO, CFO, or VP-level) with active engagement, not just nominal support. Executives must participate in FinOps reviews and allocate budget for tooling and headcount.

AI Reasoning

Final maturity assessment: Partial. Evidence-check resolved the scanner score from 2 to 1 after a targeted rescan. Verifier status: weak. The evidence supports active CTO/CFO engagement through quarterly commitment and architecture cost reviews and escalation of unresolved risks to the CTO/CFO steering group. It does not support executive allocation of dedicated budget for FinOps tooling, training, or headcount, and it does not establish cloud cost efficiency as a standing agenda item in broader executive leadership meetings. The scanner's count of 2 is not supported by its own reasoning or the raw material.

Evidence
  • “Quarterly: commitment strategy and architecture cost review with CTO/CFO steering group.” — Cadence · Process
  • “The council meets monthly, reviews cloud budget variance, and reports unresolved cost risks to the CTO and CFO steering group.” — FinOps Council · Process
E4

Cross-Functional Collaboration (Finance-Engineering-Business)

OK

FinOps requires active collaboration between finance, engineering, and business stakeholders. These groups must share a common language, common tools, and common goals around cloud investment.

AI Reasoning

Crit 1: Found. Finance, Platform Engineering, Security, and Product Operations all collaborate on the FinOps Council with shared governance; a common allocation model and shared dashboard tooling is described. Crit 2: Found. A structured monthly budget variance review brings Finance and engineering leads together; a quarterly steering with CTO/CFO also exists; weekly optimization backlog reviews further embed cadence. Crit 3: Found. The RACI explicitly distributes ownership across Finance (budgets), FinOps (allocation/optimization), Platform Engineering (guardrails), and product teams (cost decisions) — this is a shared accountability model, not an 'engineering buys, finance pays' model. Total: 3.

Evidence
  • “The FinOps Council owns cloud financial governance with named representatives from Finance, Platform Engineering, Security, and Product Operations.” — Cloud Financial Governance Policy · Process
  • “Monthly: budget variance review with Finance and engineering leads.” — Cadence · Process
  • “The RACI states that Finance owns budgets, FinOps owns allocation and optimization process, Platform Engineering owns guardrails, and product teams own cost decisions for their services.” — Operating model · Accountability
  • “Cloud costs are allocated weekly to product teams using tags, account hierarchy, and shared platform allocation rules. Team leads receive showback views for product, environment, and owner_team.” — Showback model · Financial-Integration
E5

Continuous Improvement & Benchmarking

NOK

The organization must treat FinOps as a continuous improvement discipline, not a one-time project. Internal benchmarking (month-over-month, team-vs-team) and external benchmarking (industry peers) must drive ongoing maturity.

AI Reasoning

Final maturity assessment: NOK. Evidence-check resolved the scanner score from 1 to 0 after a targeted rescan. Verifier status: weak. The quoted source supports ongoing optimization activity, weekly recommendation generation, and waste tracking against a target. However, E5 specifically requires continuous improvement with retrospectives and maturity assessments, internal benchmarking, and external/peer benchmarking. None of those are explicitly evidenced. The scanner correctly noted the gaps, but counting one criterion as met is too strong based on the available material.

F · GenAI & AI Cost Management

F1

Token & Model Spend Visibility

NOK

The organization must be able to see AI token, API, model, and inference spend by application, team, use case, customer, and environment. AI consumption must be visible as an operational cost surface, not only as a provider invoice line.

AI Reasoning

Crit 1: No evidence that the organization tracks token, API, model, or inference spend by application, use case, team, customer, or environment. Documents cover general cloud billing visibility only. Crit 2: No mention of provider AI costs, model endpoints, retry costs, cache costs, batch inference, or real-time inference cost drivers. Crit 3: No AI-specific dashboards or alerts for token cost trends or anomalies described. Total: 0.

F2

AI Cost Allocation & Unit Economics

NOK

The organization must map GenAI costs to products, workflows, customers, transactions, or business outcomes. AI costs should be connected to unit economics so teams can judge whether usage is economically justified.

AI Reasoning

Crit 1: No evidence of LLM or AI platform costs allocated to products, workflows, customers, teams, or business units. Crit 2: No shared AI platform cost allocation model described. Crit 3: No AI-specific unit metrics such as cost per AI task, per AI interaction, or per model call. General unit economics (checkout, loyalty API) exist but are not AI-related. Total: 0.

F3

Model Routing & Prompt Efficiency

NOK

The organization must optimize GenAI cost through model routing, prompt design, context management, caching, batching, retry controls, and cost-quality tradeoff measurement. Premium models should be used intentionally where quality, risk, or reasoning need justifies them.

AI Reasoning

Crit 1: No model-routing policy or model selection governance described. Crit 2: No mention of prompt length, context size, retrieval scope, caching, batching, retry behavior, or agent loops. Crit 3: No cost-quality-latency tradeoff measurement for AI models described. Total: 0.

F4

AI Budgeting, Forecasting & Guardrails

NOK

The organization must govern AI consumption with budgets, quotas, alerts, approvals, anomaly detection, and forecasting. AI workloads need guardrails against runaway spend from high-volume usage, retries, context growth, or experimental systems becoming production traffic.

AI Reasoning

Crit 1: No AI budgets, quotas, alerts, or approval thresholds for AI applications, model tiers, or AI environments described. Crit 2: No AI-specific spend forecasting or anomaly detection for token volume, retries, or context growth described. Crit 3: No production AI guardrails such as rate limits, spend limits, or model-access controls described. General infrastructure guardrails (autoscaling cost ceilings) exist but do not apply to AI workloads. Total: 0.

F5

AI Value Realization & Governance

NOK

The organization must compare AI consumption cost against business value, productivity gain, quality improvement, risk reduction, or revenue impact. AI governance should include cost ownership and decisions about whether AI usage is worth continuing.

AI Reasoning

Crit 1: No comparison of AI costs against business value such as productivity gain, quality improvement, risk reduction, or revenue impact. Crit 2: No AI cost governance structure involving product, engineering, finance, risk, or business stakeholders. Crit 3: No process for reviewing, redesigning, downgrading, or retiring low-value AI use cases. Total: 0.

Forensic Audit: Anti-Patterns

A · Cost Visibility & Allocation

A1

Tag Sprawl & Missing Tags

Tested absent

Cloud resources lack consistent tagging or have inconsistent, redundant tag taxonomies. Cost allocation is impossible or unreliable. Untagged spend is treated as acceptable overhead rather than a governance failure.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The raw material shows the opposite of tag sprawl: mandatory tag keys, daily defect reporting, remediation SLA, exception expiry, and policy-as-code checks for new Terraform modules. No evidence supports significant untagged spend, ad-hoc taxonomy, or acceptance of untagged spend as overhead. Coverage interpretation: Relevant tagging governance and remediation coverage is present in src-001-c001 and src-004-c001, so this anti-pattern would likely have been revealed if present.

A2

Black Box Cloud Spend

Tested absent

Cloud costs are visible only to a small group (typically finance or a single cloud admin). Engineering teams cannot see the cost of the resources they provision. Cost is discovered only during monthly invoice shock.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source contradicts black-box cloud spend: engineering has resource-level dashboard views and drill-downs, dashboards refresh every 12 hours, and alerts route to owning teams. No evidence suggests cost visibility is restricted to finance or a single admin. Coverage interpretation: Dashboard and alert-routing evidence in src-002-c001 directly covers engineering visibility, making absence meaningful.

A3

Delayed Cost Reporting

Tested absent

Cost data is stale — reports arrive days or weeks after the spend occurs. By the time anomalies are discovered, the damage is done. No near-real-time cost signal exists.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source supports 12-hour dashboard refreshes and daily automated anomaly detection with Slack/on-call routing. There is no evidence of reporting delays over 48 hours, manual spreadsheet-based reporting, or lack of automated anomaly alerting. Coverage interpretation: Cost reporting cadence and anomaly automation are explicitly covered in src-002-c001, so the absence of delayed reporting is reasonably tested.

A4

Siloed Cost Views

Tested absent

Each cloud provider, account, or team has its own cost view with no unified perspective. Total cloud economics is unknown. Finance sees invoices; engineering sees resource metrics; nobody sees both together.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source describes a unified FinOps dashboard integrating AWS, Azure, and GCP billing exports with role-appropriate views. There is no evidence of isolated provider/team views or finance and engineering using conflicting tools or numbers. The unit-economics gap does not by itself evidence siloed cost views. Coverage interpretation: Cross-cloud dashboard integration and shared stakeholder views are explicitly covered in src-002-c001, so absence of provider/account silos is meaningfully tested.

A5

Vanity Cost Metrics

Not assessed

The organization tracks meaningless or misleading cost metrics — total spend without context, percentage discounts without baseline, or savings numbers that cannot be verified. Metrics look good on slides but drive no real optimization.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 1 to 0. Verifier status: unsupported. Adjudication: The cited sources show limited unit economics maturity and some savings reported as absolute EUR run-rate reduction, but they do not show misleading, unverifiable, presentation-only, or non-actionable vanity metrics. The optimization evidence indicates operational use through right-sizing recommendations, product-team approval/rejection, and backlog decisions. Coverage interpretation: Source coverage discusses unit economics gaps and optimization reporting, but does not provide enough targeted evidence about metric definitions, dashboard intent, or savings-verification practices to support a harmful vanity-metrics finding.

B · Rate & Usage Optimization

B1

Commitment Avoidance

Tested absent

The organization runs predominantly on on-demand pricing despite stable, predictable workloads. Fear of commitment (vendor lock-in anxiety, forecasting uncertainty) results in paying 40-70% premiums unnecessarily.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The raw material contradicts commitment avoidance: it documents active commitment management, defined target coverage, current coverage, monthly utilization review, escalation of unused commitments, committed-discount utilization tracking, and quarterly strategy review. No source evidence supports fear-driven avoidance or predominantly on-demand use for stable workloads. Coverage interpretation: Commitment management is directly covered in src-003-c001 and src-002-c001. The evidence is specific enough that a commitment-avoidance pattern would likely be visible if present.

Evidence
  • “Target commitment coverage is 72% for stable compute and database workloads. Current coverage is 68%. Utilization is reviewed monthly; unused commitments above EUR 5,000 monthly amortized value are escalated to Finance and Platform Engineering.” — Commitment management · Operational
B2

Chronic Over-Provisioning

Partial finding

Resources are systematically oversized 'just in case.' Instances run at 5-15% CPU utilization. Databases are provisioned for peak capacity that never arrives. Nobody right-sizes because nobody owns the waste.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 1 to 0. Verifier status: unsupported. Adjudication: The only harmful signal is that 7 right-sizing recommendations were deferred due to release freeze risk, while 22 were accepted and reduced monthly run-rate by EUR 46,000. The source does not show chronic over-provisioning, utilization below 20%, unvalidated peak-capacity justification, or ignored recommendations due to lack of ownership. Coverage interpretation: Relevant right-sizing coverage exists in src-003-c001 and src-004-c001, but it tends to contradict the anti-pattern by showing weekly utilization-based recommendations and product cost owner review. The scanner signal is too weak to support a partial finding.

B3

Resource Hoarding & Zombie Resources

Tested absent

Teams accumulate cloud resources they no longer use — stopped instances with attached storage, orphaned snapshots, unused IP addresses, dev environments left running permanently. Cleanup is nobody's job.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source shows non-production shutdown outside business hours, owner-approved exceptions, weekly detection of idle/orphaned resources, and waste reporting with a reduction target. This contradicts the core anti-pattern conditions of dev/test running permanently and no lifecycle management or cleanup ownership. Coverage interpretation: Waste controls are directly covered in src-003-c001. Although residual waste exists, the source includes controls that would reasonably reveal hoarding/no-cleanup behavior if present.

Evidence
  • “Non-production environments shut down outside business hours unless tagged always_on=true with owner approval. Idle load balancers, unattached disks, and orphaned snapshots are detected weekly.” — Waste controls · Automation
B4

Manual-Only Optimization

Partial finding

Cost optimization is a manual, periodic exercise — someone reviews a spreadsheet quarterly and makes recommendations. There is no automation, no continuous optimization, no integration with engineering workflows.

AI Reasoning

Final anti-pattern assessment: Partial finding. Evidence-check resolved the scanner score from 1 to 1. Verifier status: weak. The cited evidence shows some manual or incomplete automation gaps: CI runner fallback is manual, Infracost is not enforced across all product repositories, and legacy manually provisioned resources remain. However, these do not support the Manual-Only Optimization anti-pattern criteria. The source also documents weekly generated optimization recommendations, weekly backlog review, automated anomaly detection, dashboarding, and non-production shutdowns. The positive score is too strong for the harmful pattern. Coverage interpretation: There are real partial manual-process signals, but not enough to verify the manual-only anti-pattern. The environment is not shown to lack optimization automation or engineering workflow integration entirely.

Evidence
  • “Interruption handling exists for analytics batch jobs, but CI runner fallback is still manual.” — Spot and preemptible usage · Operational
  • “Infracost estimates are visible in pull requests for the platform account, but not yet enforced for all product repositories. Legacy manually provisioned resources still exist in two older analytics accounts.” — Infrastructure as Code · Automation
B5

Discount Stacking Ignorance

Not assessed

The organization does not understand how discounts interact across providers. Savings Plans applied to already-discounted workloads, conflicting commitment types, or unused negotiated rates indicate a lack of rate optimization sophistication.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: weak. The source supports commitment utilization tracking and multi-cloud commitment management, which argues against a total absence of rate optimization governance. However, it does not specifically evidence discount-interaction understanding, overlap analysis, or avoidance of double-coverage. The scanner's zero positive score is acceptable, but its claim that the source directly contradicts all discount-stacking ignorance criteria is overstated. Coverage interpretation: The available commitment evidence is relevant but not detailed enough to prove absence of discount-overlap or stacking-governance issues. Targeted evidence on overlap analysis and discount interaction governance would be needed.

C · Governance & Policy

C1

Shadow IT & Unmanaged Cloud Accounts

Tested absent

Teams or individuals create cloud accounts outside the governed organizational structure. Spend occurs on personal credit cards, department accounts, or unlinked accounts. Total cloud exposure is unknown.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. No source chunk evidences personal-card spend, unlinked accounts, unknown cloud exposure, or cloud accounts outside the governed structure. The noted legacy manually provisioned resources and development accounts without Slack routing are governance gaps, not direct shadow IT evidence. Coverage interpretation: The packet includes relevant governance, tagging, billing-dashboard, account-dimension monitoring, and known account limitation material. These sources would reasonably surface unmanaged account exposure if documented in the available evidence.

C2

Budget Blowout Tolerance

Tested absent

Cloud budgets are set but not enforced. Overruns are discovered after the fact with no consequences. Forecasts are fiction. The budget process is a compliance exercise disconnected from operational reality.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source material shows active budget governance: monthly budget-to-actual comparison, variance flagging above 10%, forecasting inputs, monthly reviews, and corrective actions. This contradicts budget blowout tolerance. Coverage interpretation: Budgeting and forecasting are directly covered in src-001-c001 and src-004-c001, with enough process detail to assess whether overruns are ignored or disconnected from operations.

C3

FinOps Theater (Process Without Teeth)

Tested absent

The organization has FinOps titles, meetings, and dashboards but no actual optimization outcomes. FinOps is a checkbox for management reporting, not an operational discipline. Meetings happen but nothing changes.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The source material contains measurable FinOps outcomes, including accepted right-sizing recommendations, EUR 46,000 monthly run-rate reduction, and an anomaly remediation with EUR 18,400 estimated monthly avoided cost. This contradicts the FinOps Theater anti-pattern. Coverage interpretation: The packet includes operating model, cadence, optimization review, and anomaly remediation evidence, which is relevant and sufficient to test whether FinOps activities produce measurable outcomes.

C4

Vendor Lock-In Blindness

Partial finding

The organization makes major cloud commitments (multi-year EDPs, proprietary service adoption) without evaluating exit costs, portability, or competitive alternatives. Lock-in is accepted by default rather than managed deliberately.

AI Reasoning

Final anti-pattern assessment: Partial finding. Evidence-check resolved the scanner score from 1 to 1. Verifier status: weak. The cited quote is real and supports a partial portability-governance gap: exit-cost analysis for proprietary managed services is only performed for tier-1 systems. However, the scanner's inference that this means major commitments below tier-1 are made without exit analysis is stronger than the raw text. Competitive benchmarking is explicitly present. Coverage interpretation: There is a real but limited harmful-pattern signal around incomplete exit-cost analysis. The same source also shows vendor pricing benchmarking, so full vendor lock-in blindness is not established.

Evidence
  • “Vendor pricing is benchmarked during annual renewal, but exit-cost analysis for proprietary managed services is only performed for tier-1 systems.” — Compliance and vendor management · Process
C5

Compliance as Cost Afterthought

Tested absent

Regulatory and compliance requirements create cost surprises because they are not integrated into architecture planning. Data residency, encryption mandates, and audit requirements add unplanned spend.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source material shows data residency requirements reviewed during architecture design and compliance-driven costs such as logging, encryption, and regional replication tracked separately. No source evidence shows compliance cost surprises or compliance treated reactively. Coverage interpretation: Compliance cost handling is directly covered in src-004-c001, and architecture standards also include data-retention cost impact. This is enough to test the stated anti-pattern in the available evidence.

D · Architecture & Engineering

D1

Lift-and-Shift Without Optimization

Partial finding

Workloads are migrated to the cloud with their on-premises sizing intact. Virtual machines mirror physical server specs. No cloud-native optimization occurs post-migration. The cloud becomes an expensive colocation facility.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: weak. No source evidence shows workloads migrated with on-premises sizing intact, VMs matching physical server specs, or cloud being treated as colocation. However, the packet is mostly current FinOps/engineering governance and optimization material, not migration-history evidence, so the scanner’s strong claim that absence is fully verified is too strong. Coverage interpretation: The sources contain relevant current optimization signals, including right-sizing and architecture cost reviews, but they are largely silent on original migration approach and inherited on-prem sizing. Absence of this anti-pattern is therefore not fully testable from this packet.

D2

Cost-Blind Architecture

Tested absent

Architecture decisions are made purely on technical merit without cost modeling. Teams choose services, instance types, and configurations based on features or familiarity, not cost-performance tradeoffs.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The source directly contradicts cost-blind architecture: architecture review templates require estimated monthly cost, scaling cost curves, data-retention cost impact, and cost-performance tradeoff sections. Infracost is also visible in some pull requests, although not enforced everywhere. No harmful cost-blind decision pattern is evidenced. Coverage interpretation: Architecture governance and deployment-cost evidence are directly covered in src-004-c001, and the documented requirements are the opposite of the anti-pattern criteria.

Evidence
  • “Architecture review templates require estimated monthly cost, scaling cost curve, data-retention cost impact, and a cost-performance tradeoff section.” — Architecture standards · Policy
D3

Scaling Without Limits

Partial finding

Autoscaling is configured without cost ceilings. A traffic spike or runaway process can scale resources to unlimited cost. There are no circuit breakers between demand signals and infrastructure provisioning.

AI Reasoning

Final anti-pattern assessment: Partial finding. Evidence-check resolved the scanner score from 1 to 1. Verifier status: weak. The checkout-api incident in src-002-c001 supports one harmful-pattern criterion: compute spend increased 31% after a misconfigured autoscaling threshold. The broader anti-pattern is not fully supported because src-004-c001 states autoscaling policies must define maximum capacity, target utilization, and cost ceiling guardrails, indicating governance exists or is being remediated. Coverage interpretation: There is a real but limited harmful signal: one documented autoscaling-related cost spike. The available source also shows corrective action and policy guardrails, so the anti-pattern is only partially evidenced rather than broadly embedded.

Evidence
  • “On 2026-03-18, checkout-api compute spend increased 31% after a misconfigured autoscaling threshold. The alert was routed to Platform Checkout and FinOps. The team corrected the scaling policy, created a follow-up architecture review.” — Recent anomaly example · Operational
  • “Autoscaling policies must define maximum capacity, target utilization, and cost ceiling guardrails.” — Architecture standards · Policy
D4

Single-Cloud Tunnel Vision

Tested absent

The organization is locked into a single cloud provider without evaluating whether specific workloads would be more cost-effective elsewhere. Provider loyalty overrides economic rationality. Competitive pricing pressure is absent.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The evidence contradicts single-cloud tunnel vision: the dashboard integrates AWS, Azure, and GCP billing exports, commitments are managed across all three providers, and vendor pricing is benchmarked during annual renewal. The limited exit-cost analysis for proprietary services is a vendor-risk gap, not evidence of single-cloud provider loyalty. Coverage interpretation: The packet has direct multi-cloud cost-management and vendor-benchmarking evidence, which would reasonably reveal this anti-pattern if the organization were single-cloud or avoided provider comparison.

Evidence
  • “The FinOps dashboard refreshes every 12 hours and integrates AWS, Azure, and GCP billing exports.” — Dashboards · Operational
  • “Vendor pricing is benchmarked during annual renewal, but exit-cost analysis for proprietary managed services is only performed for tier-1 systems.” — Compliance and vendor management · Process
D5

Monolith Tax

Partial finding

Large monolithic applications prevent granular cost allocation and optimization. The entire monolith must be scaled together even if only one component is under load. Decomposition is deferred indefinitely.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: weak. No source evidence shows monolithic applications preventing granular allocation, whole-application scaling, or deferred decomposition. However, the source packet does not materially discuss application architecture or decomposition, so the scanner’s claim that absence is adequately verified is too strong. Coverage interpretation: The sources include service/product cost visibility, but they are not detailed application-architecture documents. They do not provide enough coverage to test whether monolith-related scaling or allocation constraints exist.

E · Culture & Organization

E1

Cost is IT's Problem

Tested absent

Cloud cost management is viewed as solely an IT or infrastructure responsibility. Business stakeholders who drive demand feel no accountability for cost. Engineers who provision resources never see the bill.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The raw source describes the opposite of 'Cost is IT's Problem': product domains appoint engineering cost owners, product teams own cost decisions for their services, team leads receive showback, and anomaly alerts route to owning teams. No harmful-pattern criterion is evidenced. Coverage interpretation: Coverage is directly relevant to ownership and accountability. The operating model, showback policy, and anomaly routing would reasonably reveal if cost were isolated to IT; instead they show distributed accountability across Finance, FinOps, Platform Engineering, and product teams.

Evidence
  • “Each product domain appoints an engineering cost owner... product teams own cost decisions for their services.” — Operating model · Accountability
E2

Blame-Based Cost Management

Tested absent

Cost overruns trigger blame and punishment rather than systemic analysis. Teams hide cost problems instead of surfacing them. Cost discussions are adversarial between finance and engineering rather than collaborative.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source does not show blame, punishment, hidden cost problems, or adversarial finance-engineering interactions. The anomaly example shows routed alerting, correction by the owning team, follow-up architecture review, and recorded avoided cost. Corrective actions are also captured in a FinOps action log. Coverage interpretation: Coverage includes an anomaly runbook, a recent anomaly response, budget variance process, and cross-functional review cadence. These are relevant enough to test for blame-based behavior, and the described behavior is collaborative rather than punitive.

Evidence
  • “The alert was routed to Platform Checkout and FinOps. The team corrected the scaling policy, created a follow-up architecture review, and recorded estimated monthly avoided cost.” — Recent anomaly example · Process
  • “Corrective actions are captured in the monthly FinOps action log.” — Budget and forecast · Process
E3

FinOps Lip Service

Tested absent

Leadership talks about FinOps but does not invest in it. No dedicated team, no tooling budget, no training. FinOps is an additional duty for someone who already has a full-time job. Optimization is expected to happen organically.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source supports a staffed FinOps function, a federated cost-owner model, recurring reviews, dashboards/anomaly tooling, optimization backlog processes, and executive steering engagement. That contradicts the main FinOps Lip Service indicators of no dedicated team, no structured program, and optimization expected to happen organically. Explicit training budget is not shown, but no harmful criterion is positively evidenced. Coverage interpretation: Coverage includes operating model, headcount, tooling, governance cadence, and optimization processes. These materials would reasonably expose whether FinOps were merely nominal; instead they show operational implementation.

Evidence
  • “The central FinOps team has one FinOps lead, one cloud economist, and one platform analyst.” — Operating model · Operational
  • “Quarterly: commitment strategy and architecture cost review with CTO/CFO steering group.” — Cadence · Process
E4

Finance-Engineering Wall

Tested absent

Finance and engineering operate as separate fiefdoms with different tools, languages, and objectives around cloud spend. Finance sees cost centers; engineering sees resource metrics. Translation between them is manual and lossy.

AI Reasoning

Final anti-pattern assessment: Tested absent. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The source shows shared cross-functional forums and tooling rather than a finance-engineering wall: the FinOps Council includes Finance, Platform Engineering, Security, and Product Operations; monthly reviews bring Finance and engineering leads together; and the dashboard includes executive, finance, and engineering views. No harmful-pattern criterion is evidenced. Coverage interpretation: Coverage is directly relevant to finance-engineering collaboration, tooling, and governance forums. Although individual financial/technical literacy is not assessed, the available source shows structures that bridge rather than separate finance and engineering.

Evidence
  • “The FinOps Council owns cloud financial governance with named representatives from Finance, Platform Engineering, Security, and Product Operations. The council meets monthly, reviews cloud budget variance.” — Cloud Financial Governance Policy · Process
  • “The FinOps dashboard refreshes every 12 hours and integrates AWS, Azure, and GCP billing exports. It contains executive trend views, finance budget views, and engineering resource-level views.” — Dashboards · Operational
E5

Static Maturity Assumption

Not assessed

The organization treats FinOps maturity as a destination rather than a continuous journey. After initial optimization efforts, the discipline atrophies. No regular maturity assessment, no improvement targets, no benchmarking.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0 after a targeted rescan. Verifier status: supported. The source does not positively evidence a static maturity assumption. It shows ongoing optimization activity, weekly recommendations, monthly/quarterly review cadences, waste targets, and measured run-rate reduction. However, the source is also silent on formal FinOps maturity assessments and external maturity benchmarking, so absence of the anti-pattern is not fully testable. Coverage interpretation: The available material meaningfully contradicts stalled optimization momentum, but it does not fully cover maturity-assessment or external benchmarking practices. Therefore the harmful pattern is not evidenced, but full absence cannot be confirmed.

F · GenAI & AI Cost Management

F1

Invisible Token Spend

Not assessed

AI/API usage exists, but token and model costs are not visible by owner, use case, application, or system. GenAI cost appears only as a provider invoice or platform total, preventing accountability.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The scanner correctly does not score Invisible Token Spend as present. The source does not establish that AI/API/token/model spend exists at all, so invisible AI spend cannot be confirmed from these documents. Coverage interpretation: The documents cover general cloud FinOps governance, but are silent on AI consumption. Because there is no evidence of existing AI spend, absence of the anti-pattern is not tested; it is unknown/not assessable.

F2

Playground-to-Production Cost Drift

Not assessed

Experiments, pilots, or playground usage become production AI consumption without cost governance. Sandbox keys, prototypes, or ad hoc integrations accumulate persistent spend without lifecycle review.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. No AI experiments, playgrounds, sandbox AI keys, prototypes, or AI production transitions are described. The scanner correctly leaves the anti-pattern unscored rather than asserting absence. Coverage interpretation: The supplied material is silent on AI lifecycle or GenAI workload promotion. This does not prove the anti-pattern is absent; it only makes it not assessable from the provided evidence.

F3

Premium Model Overuse

Tested absent

Expensive models are used where cheaper models, caching, batching, smaller context, or simpler workflows would meet the need. Model choice is driven by convenience or prestige rather than measured cost-quality fit.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. There are no references to AI models, premium model usage, cheaper model alternatives, AI caching/batching, prompt reduction, retries, or agent loops. Premium Model Overuse is not evidenced. Coverage interpretation: The source is silent on AI model usage and model-selection practices, so the anti-pattern cannot be tested absent.

F4

Unbounded Context Growth

Tested absent

RAG, prompt, memory, or conversation context grows without cost-quality control. The organization pays for excessive tokens because retrieval scope, history length, documents, or prompt structures are not governed.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. The documents contain no RAG, prompt, memory, context-window, conversation-history, retrieval-scope, cache-efficiency, or token-budget evidence. Unbounded Context Growth is not evidenced. Coverage interpretation: The source does not cover AI context or retrieval architecture. Absence is therefore unknown rather than tested absent.

F5

AI Value Theater

Tested absent

AI usage is celebrated while cost per outcome is unknown. The organization tracks usage, pilots, or activity volume but cannot show whether AI spend improves productivity, quality, revenue, risk, or customer outcomes.

AI Reasoning

Final anti-pattern assessment: Not assessed. Evidence-check resolved the scanner score from 0 to 0. Verifier status: supported. No AI adoption metrics, AI pilot reporting, AI activity dashboards, AI ROI claims, or continued AI use cases with unclear value are present in the source. AI Value Theater is not evidenced. Coverage interpretation: The documents are silent on AI adoption/value reporting, so the anti-pattern cannot be meaningfully tested absent.

Quality & Strategy Hygiene Appendix

Quality Gate detail is retained here for traceability. WARN-level strategy hygiene notes do not invalidate the assessment score.

Reviewer Summary · gpt-5.5

The assessment is not blocked, but the WARN result means some scores and roadmap items were cleaned up because the source evidence did not consistently support them. Strategy hygiene notes were retained for traceability; they do not invalidate the score. Treat the maturity reading as usable, with caution around weakly grounded zero scores and remaining unsupported claims.

Sanitized strategy items
Evidence-check adjustments
Strategy hygiene notes
Remaining warnings
Fact-check trajectory
Remaining fact-check notes

Source Registry & Context Packets

This snapshot shows how parsed source material was chunked, sampled for DLP review, and routed into A-F context packets before model audit. Packetization controls attention, not truth: findings still require verified source evidence.

MetricValue
Source documents4
Parsed chunks4
DLP review chunks4
High-risk DLP hits0
Caution DLP hits0
PacketIncluded chunksCandidate chunksCoverageCharacters
ACost Visibility & Allocation 4 4 OK 8,319
BRate & Usage Optimization 4 4 OK 8,317
CGovernance & Policy 4 4 OK 8,301
DArchitecture & Engineering 4 4 OK 8,319
ECulture & Organization 4 4 OK 8,315
FGenAI & AI Cost Management 4 4 Weak coverage 8,320

RunTrace Provenance

RunTrace is a client-side provenance artifact. It records source/chunk references, hashes, model-stage metadata, evidence paths, score paths, tactic paths, and Quality Gate decisions without embedding full raw source documents or full prompts.

Run ID20260723064613-wh3a3n4zq
Sources4
Chunks4
Model Stages43
Evidence Paths74
Score Paths60
Tactic Paths26
DLP Chunks4
GateWARN
Trace boundaries