
The email arrives at 09:14 and it is three sentences long. Something is wrong, their team is stuck, and they would like to know what you are going to do about it. What you owe them in that moment is not one thing. It is four different obligations that arrived in the same email and that founders answer as though they were a single question — what the contract says you owe, what the law requires you to tell them, what the relationship needs in the next hour, and what you owe the next customer by writing down what happened. They have different deadlines, different currencies and different people who decide whether you met them. Answering the wrong one first is the characteristic mistake, and it is expensive in a way that does not show up until renewal.
Start with the contract, because it almost certainly says less than you think and possibly nothing at all. Most first serious customer agreements have no service level in them. The ones that do usually contain a clause somebody copied from a cloud provider, and it is worth reading what that clause actually does in its original home. Amazon's compute agreement promises to “use commercially reasonable efforts to make Amazon EC2 available for each AWS region with a Monthly Uptime Percentage of at least 99.99%”, and if it misses, the remedy is a percentage discount: 10% for a month under 99.99% but at or above 99.0%, 30% under 99.0%, 100% under 95.0%. Then come the parts nobody copies across deliberately. “To receive a Service Credit, you must submit a claim by opening a case in the AWS Support Center”, and the request “must be received by us by the end of the second billing cycle after which the incident occurred”. “Service Credits will not entitle you to any refund or other payment from AWS.” The credit is only issued at all if it comes to “greater than one dollar ($1 USD)”. And the clause that does the real work: “Unless otherwise provided in the Agreement, this SLA sets forth your sole and exclusive remedies, and AWS’ sole and exclusive obligations, for any unavailability, non-performance, or other failure by us to provide Amazon EC2” (Amazon Web Services, Amazon Compute Service Level Agreement, last updated 25 May 2022).
Read that as a whole and it is not a promise. It is a liability cap wearing the costume of a guarantee: the maximum you can recover is a discount on what you already paid, you have to go and ask for it inside a deadline, and asking for it is the only thing you are allowed to do. There is nothing wrong with a cap — an infrastructure business that carried uncapped consequential losses would not exist — and a founder is entitled to one too. What is wrong is pasting the clause into your own agreement and then believing that the number it produces is what you owe. The cap is the ceiling on what your customer can demand. It has never been the floor on what you should offer, and the gap between those two is where the account is actually won or lost.
The second obligation is legal, and the thing to understand about it is that it does not trigger on “broke”. Founders get this wrong in both directions — they panic about an outage that triggers nothing, and they stay quiet about a leak that triggers a statute with a penalty attached. In Virginia the trigger is narrow and precisely drawn. A reportable breach means “the unauthorized access and acquisition of unencrypted and unredacted computerized data that compromises the security or confidentiality of personal information... and that causes, or the individual or entity reasonably believes has caused, or will cause, identity theft or other fraud to any resident of the Commonwealth”, and “personal information” is defined as a name in combination with a social security number, a driver's license or state identification card number, a financial account or card number together with the code that would let someone use it, a passport number, or a military identification number (Code of Virginia, § 18.2-186.6, Breach of personal information notification, retrieved 28 September 2026).
Three consequences fall out of that definition and every one of them surprises somebody. Access is not enough — the statute says access AND acquisition. Encryption is a defence, which is why the question “was it encrypted at rest, and did whoever got in also have the key?” is the first technical question of the morning rather than the fifth. And a database of email addresses and hashed passwords, which is what most product breaches actually consist of, is not personal information under this section at all. That is not permission to say nothing; it means the duty you are under is the commercial one rather than the statutory one, and you should know which of the two you are discharging while you draft the message.
When it does trigger, the statute writes your notice for you, and it is worth pre-drafting because you will not want to compose it at 03:00. Notice goes to “the Office of the Attorney General and any affected resident of the Commonwealth without unreasonable delay”, and it must describe “(1) The incident in general terms; (2) The type of personal information that was subject to the unauthorized access and acquisition; (3) The general acts of the individual or entity to protect the personal information from further unauthorized access; (4) A telephone number that the person may call for further information and assistance, if one exists; and (5) Advice that directs the person to remain vigilant by reviewing account statements and monitoring free credit reports.” Notify more than 1,000 people at once and the nationwide consumer reporting agencies have to be told as well. The Attorney General may seek “a civil penalty not to exceed $150,000 per breach of the security of the system or a series of breaches of a similar nature that are discovered in a single investigation” (same section, subsections B, E and I).
One structural point that founders outside Virginia should take from this rather than the specific numbers: the obligation follows the resident, not your office. A company in Glen Allen with customers in four states is potentially subject to four statutes with four definitions and four clocks, and the fastest of them governs. That is an argument for knowing where your users actually are before anything happens, which is a data-inventory question, not an incident question. It is the same inventory that answers what you are renting and what you own — do you own your product, or are you renting it? sets out how to build it — and if a processor of yours is the one that leaked, subsection D puts a duty on them to tell you, which is only useful if your contract with them says the same thing.
The third obligation is the one that decides whether you keep the account, and no statute mentions it. NIST's incident response guidance separates communication into four kinds, and the distinction is the most useful thing a founder can take from the document. Incident coordination is “communicating current and planned incident response activities for a particular incident among the internal and external parties with incident response roles and responsibilities.” Incident notification is “formally informing affected customers, employees, partners, regulators, or others about a data breach or other incident.” Public communication is “communicating to the public about the status of a particular incident, such as responding to media inquiries.” Incident information sharing is voluntary threat sharing with others. The guidance then adds the recommendation that costs nothing in peacetime and cannot be executed under pressure: “Organizations should have mechanisms in place in advance to coordinate with affected parties about incidents when needed” (National Institute of Standards and Technology, Incident Response Recommendations and Considerations for Cybersecurity Risk Management, SP 800-61r3, April 2025, which supersedes the 2012 Revision 2).
Note that phrase: or other incident. Notification is not a breach word. The category covers the outage, the corrupted import, the two weeks during which a report was silently wrong — and that last one is the case founders handle worst, because nobody is shouting and it is possible to fix it quietly and say nothing. Do not. A customer who discovers in March that a number they used in January was wrong and that you knew in February has learned something about you that no amount of subsequent uptime corrects.
The first message goes out before you know the cause, and getting that order right is most of the skill. Timing beats completeness, and the reason is that your customer is not waiting for an explanation — they are waiting to find out whether they have to tell their own customers something. So the first message contains four things and no more: what is affected, what is confirmed not affected, what you are doing right now, and the time of the next update. It does not contain the cause, because your first theory of the cause is wrong about a third of the time and the correction is worse than the silence would have been. It does not contain a restoration estimate, for the same reason. The only commitment you can make in hour one that you are certain to keep is the time of the next message, so make that one and keep it, including the update that says nothing has changed. Frequency is the signal. Founders under-communicate during the long middle of an incident precisely when the customer's anxiety is compounding.
Then the question of money, where there are three currencies and they are not interchangeable. A credit is the cheapest and the least meaningful: it is a discount on future business with the party who just let you down, and offering it as the opening move reads as a negotiation rather than an apology. A refund is real money and it says you do not think you earned the period in question. Remediation is the expensive one and the only one that maps to an actual loss — paying for the hours their team spent working around you, rebuilding the data your system corrupted, covering the cost of the notice they now have to send their own users. The rule is simple to state and uncomfortable to follow: offer the currency that matches the loss you caused, not the one that matches the size of their invoice. If a four-hour outage cost them a day of six people's work, a 10% credit on a monthly subscription is arithmetic they can do faster than you can.
Data loss is a different obligation from downtime and deserves its own sentence in the message. Restoring from a backup is not a return to normal; it is a decision to discard everything between the snapshot and the failure, and the work in that window belonged to the customer. Tell them the exact boundary — this timestamp to this timestamp — and tell them what falls inside it, because they may be able to reconstruct it and will certainly resent finding out by accident. This is also where an untested restore becomes a commercial problem rather than an engineering one, and where the drills that prove the system can be recovered by somebody other than its author earn their afternoon: how to tell if your system is production-ready sets out the shortest version of them.
The write-up is owed to them too, and there is a specific error to avoid about when to send it. The temptation is to wait until the analysis is complete so that the document is clean. NIST's guidance is explicit on the other side: “The lessons learned during incident response should often be shared as soon as they are identified, not delayed until after recovery concludes.” Send the interim version. It should say what happened in plain language, what the impact was including the parts you did not notice first, what has already changed, and what will change with a date beside it. Two things keep it useful. It names no individual — a document that identifies the engineer who ran the migration is a document about your company's culture and not about the incident, and it guarantees that the next one gets reported late. And it does not promise a process where a control belongs; a customer who reads “we have added a review step” hears a human being promising to be more careful, which is the least durable fix available.
The four things not to do are all things that feel reasonable at the time. Do not go quiet while you investigate; the investigation is not the story, the silence is. Do not disclose partially and get corrected later, because a second disclosure that enlarges the first is treated as the discovery of a cover-up even when it is honest sequencing — which is why the first message should be narrow and certain rather than broad and provisional. Do not blame a vendor, however true it is: your customer bought from you, they have no contract with your provider, and the statement “our cloud provider had an outage” tells them only that you chose the provider and had no fallback. And do not offer compensation before you know the scope. A generous gesture on Tuesday that turns out on Thursday to have been a fraction of the damage reads as an attempt to settle cheaply while you still knew more than they did.
Almost none of this can be composed under pressure, and the preparation is about one page long. Who decides the product is broken, by name, and who can declare it over. Where the customer-facing update is published and who has the credentials to publish it — a status page nobody can reach during an incident is a common and entirely self-inflicted failure. The template for the first message, with the four fields already laid out. And the list of who must be told and inside what window, with your customers' states written down beside the addresses. Five jobs have to carry a name on the day a system goes live and most launches cover all five with whoever built it, which is the same structural gap seen from a different angle — who owns an AI system after it ships names them. If that person has just handed in their notice, the week your only engineer resigns is when the incident plan stops being paperwork.
The reason to take all of this seriously at three customers rather than at thirty is that the first serious incident is the one that sets your reputation with a customer who has not yet decided about you. Early users are more forgiving of failure than founders expect and far less forgiving of being managed. The product breaking is evidence that you are building something real; the handling of it is the only evidence they will ever get about how you behave when something is at stake, and they are reading it that way whether or not you intend them to. Most of what breaks on contact with real users is predictable and well catalogued — what breaks when an AI-built prototype meets real users lists six of them, and the first ninety days after launch covers the support load nobody staffs for.
So write the page now, while nothing is wrong. If you want the failure modes named before a customer finds them, the free production-readiness diagnostic does that from the outside, and where the exposure is personal data rather than uptime, that is the territory AI security and compliance services exists to cover. A startup software development company worth engaging will ask what your incident process is during the sales conversation rather than after the first outage — and if nobody has asked you that yet, the honest reading is that nobody has been thinking about the part of the work that starts after launch.
Related: Startup software development company
Find this useful? Tell Google to show you more of it.
