FIELD NOTE · SYSTEMS MAINTENANCE

OPEN ITEMS / 07

The interest
on unfinished
work.

A quick fix buys today. The bill arrives later as slower changes, brittle handoffs, vendor surprises and failures nobody can explain.

Read the fault map
MAINTENANCE LEDGER / SAMPLE VIEWRISK × FREQUENCY
ITEMWHEN IT HURTSOWNERSTATE
The one-off exportEvery FridayIan?HOT

Manual data cleaning sits between a source system and the report. The risk is not the spreadsheet; it is the missing definition of “current.”

The expired keyWhen it expiresVendorWATCH

A service account works until the day it does not. The handover question is who receives the alert and who can rotate the credential safely.

The quiet queueWhen demand risesUnknownHOT

Requests wait in an inbox with no age, owner or escalation. A queue that looks calm can be a system with no measurement.

The abandoned pathWhen a user retriesFormer devCOLD

An error path was left for later. Nobody has tested it since the last integration changed.

Illustrative ledger. The point is to name where interest can accumulate; these are not XT3 client findings.

IAN DIKHTIAR / RESEARCHKEEP THE SYSTEM OWNABLE7 sources ↗

01 / The fault map

Debt is not
ugly code.

It is a decision whose cost is deferred. Sometimes the decision is right. The failure is forgetting that the decision has a maturity date, an owner and a price that changes when the system changes.

01Shortcut

Ship the workaround to meet a real deadline.

Benefit now
02Exposure

The system changes, but the workaround is not re-evaluated.

Context shifts
03Interest

Every future change requires extra discovery, checking or repair.

Work compounds
04Choice

Pay deliberately, accept the risk, or replace the path.

Owner decides
THE USEFUL METAPHOR

Principal is the work you postponed. Interest is the extra work you pay whenever the system asks to change.

02 / The interest

Some debt is
cheap.
Some follows
every change.

The debt metaphor helps when it points to a decision: an intentional shortcut can be sensible if the benefit is real and repayment is visible.

SEI’s treatment of technical debt emphasizes that debt is about the consequences of decisions and the system’s future evolution. It recommends making debt explicit and managing it alongside other work. [1] An empirical study of eight software teams found both reactive strategies—fix it when it hurts—and systematic strategies that identify, measure and monitor debt. [2]

A separate study of 10 open-source PHP projects found that modules with more measured debt were associated with more frequent and effortful corrective maintenance in that sample. [3] The result does not give an organization a universal interest rate. It does show why “we will remember that later” is not a measurement strategy.

Decision rule

Accept a shortcut only when its benefit is named, its owner is known and its repayment trigger will appear in the planning system.

LOW INTERESTRarely touched

Document the boundary. Revisit when the workflow or dependency changes.

HIGH INTERESTEvery release

Prioritize the repair by user impact, failure exposure and effort to reverse.

Qualitative continuum only. The bars are not a debt score or measured interest rate.

03 / The reliability bill

Reliability is
a product
promise.

A system is reliable enough only relative to what people need from it. The right question is not “is it up?” but “what behavior can a user count on, under which conditions, and what happens when we miss?”

Google’s SRE guidance starts with service-level indicators and objectives: measure behavior that matters, set a target and use the gap between the target and reality to guide action. It describes error budgets as a way to balance reliability work with the pace of change. [5]

NIST’s Cybersecurity Framework 2.0 similarly offers outcomes for understanding, assessing, prioritizing and communicating risk across organizations of different sizes and maturity. It is deliberately flexible; the framework does not choose the target for you. [6]

SLIWhat happened?

Successful requests, fresh data, completed handoffs or recoverable errors.

SLOWhat do we promise?

A target and the conditions where it applies.

ACTIONWhat changes now?

Pause, repair, communicate or continue with evidence.

The examples are a decision model, not a recommended target. Set the objective with the people who bear the user and business consequences.

04 / The handover

A system that
only one person
can run is
unfinished.

Vendor choice does not transfer ownership. A platform can be healthy while the business remains unable to answer simple questions: what data enters, where it runs, which account pays, who gets the alarm and how to stop it.

QUESTIONPROOF TO KEEPIF MISSING
What does it depend on?Versions, vendors, keys, domains+

Keep an inventory of material dependencies, expiry dates and the account that controls each one. NIST’s supply-chain guidance treats dependency risk as something to identify and manage, not a surprise to discover during failure. [7]

Who can change it?Roles, access, approvals+

Write down the smallest role that can deploy, rotate a secret, edit a workflow or restore data. Test the path with a second person before the first owner is unavailable.

How do we know it works?Checks, logs, user outcome+

Keep a short runbook and a repeatable check. A green deployment is not proof that the customer’s path works end to end.

How do we stop it?Rollback, fallback, contact+

Document the safe stop and the manual fallback. Practice it on a low-risk change; recovery that exists only in a head is not recovery.

05 / The ledger

Pay the items
that keep
charging.

Do not start with a total debt score. Start with the items that touch a user, a release, a secret or a handoff often enough to matter.

Repair readiness, for a conversation
1Observe the failure2Record the exposure3Choose the smallest repair4Test the new path5Handover the knowledge

THE XT3 POINT OF VIEW

Make the system
worth inheriting.

The future owner should not have to buy the same discovery twice. I help businesses find the risky unfinished work, put a useful measure around it and repair only what changes the outcome.

Ian Dikhtiar · Hands-on fractional CTO · XT3

06 / Sources & methodology

Keep the ledger honest.

This report synthesizes public research available September 8, 2026. It makes no claim about the size of XT3 client debt, the savings from a repair or the reliability of a named vendor.

Technical-debt studies use particular organizations, projects and measurement methods. NIST and Google guidance describe management practices, not legal requirements or universal service targets. The maps, workflow examples and checklist are recommendations.

A useful first pass is a short inventory of dependencies, recurring failure modes and handoff gaps. Rank an item by user impact and frequency of exposure, then decide whether to repair, accept, isolate or replace it.

  1. [1]
    Technical Debt: From Metaphor to Theory and Practice ↗

    Philippe Kruchten, Robert Nord & Ipek Ozkaya · Carnegie Mellon Software Engineering Institute · November 2012

    SEI report that treats technical debt as a decision and management problem, recommends an explicit backlog, and distinguishes short-term advantage from future change cost.
  2. [2]
    How do software development teams manage technical debt? — An empirical study ↗

    Jesse Yli-Huumo, Andrey Maglyas & Kari Smolander · Journal of Systems and Software · October 2016

    Exploratory case study of eight teams in one large software organization, including interviews with 25 people. It identifies reactive and systematic debt-management strategies; it is not a universal maturity benchmark.
  3. [3]
    The relation between technical debt and corrective maintenance in PHP web applications ↗

    Theodoros Amanatidis, Alexander Chatzigeorgiou & Apostolos Ampatzoglou · Information and Software Technology · October 2017

    Case study of 10 open-source PHP projects. Modules with more measured technical debt showed higher corrective-maintenance probability and effort in that sample.
  4. [4]
    Guide to Enterprise Patch Management Planning: Preventive Maintenance for Technology (SP 800-40 Rev. 4) ↗

    National Institute of Standards and Technology · April 2022

    NIST guidance for planning, prioritizing and operating patch management. It calls out unsupported or end-of-life software as a reason a patch may never arrive.
  5. [5]
    Service Level Objectives ↗

    Chris Jones, John Wilkes & Niall Murphy, with Cody Smith · Google SRE · 2017 · continuously maintained online edition

    Google’s explanation of SLIs, SLOs and error budgets. It frames reliability targets as product and business decisions and uses measured outcomes to guide release and maintenance choices.
  6. [6]
    The NIST Cybersecurity Framework (CSF) 2.0 ↗

    Cherilyn Pascoe, Stephen Quinn & Karen Scarfone · NIST CSWP 29 · February 2024

    A flexible taxonomy of cybersecurity outcomes for organizations of different sizes and maturity levels. It is guidance, not a certification or a promise of compliance.
  7. [7]
    Cybersecurity Supply Chain Risk Management Practices for Systems and Organizations (SP 800-161 Rev. 1, Update 1) ↗

    National Institute of Standards and Technology · May 2022 · updated November 1, 2024

    NIST guidance for identifying, assessing and responding to cybersecurity risks in suppliers and technology dependencies. Used here for vendor ownership and handover questions, not as a procurement checklist.
Back to the beginning ↑

Notes on this report

Practical follow-ups