Start here
Does any of this sound familiar?
- Your engineering team writes postmortems that connect detection, response, contributing conditions, and follow-up actions.
- You need to evaluate whether an AI assistant can distinguish a plausible hypothesis from a supported incident finding.
- You want to preserve learning while avoiding disclosure of customer identifiers, vulnerabilities, or credentials.
Incident records show what an engineering team knew, which diagnostic branch it chose, and what evidence changed the response. Unlike polished runbooks, they preserve uncertainty, timeline, and outcome.
Training on a historical archive is not the same as building realistic evaluation tasks. A task tests an approved scenario against a rubric; a corpus contains operational history. Keep purposes, permissions, and retention distinct.
Not ready to share a single file? You don't have to.
Take the 3-question fit checkThe problem
Postmortems teach reliability, but can expose the wrong details
A report may include dependencies, alerts, failed mitigation, rollback decisions, and action owners. It may also reveal customer impact, topology, vulnerabilities, employee names, or an unfinished investigation. Internal circulation does not authorize external model access or reuse.
Evaluation requires evidence discipline. Give the model information available at the decision point, define a defensible response, and have a reviewer separate supported conclusions from speculation. Otherwise the test measures hindsight or guesswork.
A useful evaluation preserves the uncertainty that responders actually faced.
The solution
Create an incident-to-evaluation trail
Choose one model behavior to test and keep operational records separate from an external evaluation package.
DataSupply partners only with labs that meet its top 0.01% credibility standard. We help assess whether a qualified buyer may be a fit and negotiate terms that reflect the data's potential value, including exclusivity where relevant. We also help you work through diligence questions about rights, privacy, security, and compliance, then present a high-level inventory of permitted records, not the dataset. Fit is specific to each situation; no buyer or value is guaranteed.
What to inventory before any buyer conversation
- Inventory reports and classify sensitivity Record date, service, impact, owner, investigation status, and customer or security content. Exclude open investigations, legal holds, and contract-restricted reports until authorized review.
- Write tasks from approved evidence Choose one decision, such as a diagnostic check or safe rollback condition. Provide only then-known facts, define acceptable answers and unsafe actions, and link findings to evidence. Label synthetic or transformed tasks.
- Review, score, and retain the decision trail Have SRE and security or privacy reviewers approve redactions and rubric. Preserve source, version, transformations, prompt, scoring rationale, reviewer, and deletion date. Use disagreements to refine the rubric.
Set the boundaries before discussing access.
Review customer and employee confidentiality, security disclosure duties, IP, and privacy. Written terms should limit purpose, access, onward transfer, retention, deletion, derivatives, and incident response. Exclude live credentials and exploitable details.
What could make a permitted example useful?
Clear tasks and reliable labels may reduce expert evaluation effort, but report volume is not a price signal. A purpose-built task set may fit better than a corpus license; scope, rights, and validation effort matter.
A practical first step.
Map one closed, non-sensitive incident internally. Ask counsel whether it can be transformed for evaluation before sharing anything.
datasupply.ai can discuss possible fit and buyer questions without receiving your dataset. You decide whether to pursue any introduction. No buyer, license, or payment is guaranteed.
Documented example / what it proves
A published reliability practice emphasizes structured learning
Google's Site Reliability Engineering book describes blameless postmortems and recording impact, contributing causes, mitigation, and follow-up actions. It presents postmortems as a way to learn and improve systems: an example of turning operational events into decision-and-action records. Read Google Site Reliability Engineering, Postmortem Culture: Learning from Failure.
These principles can inform a rubric that separates impact, causal evidence, response, and corrective action. They do not show a postmortem license for model training or clear another company's reports for external use.
The important limit: The SRE book is guidance about reliability practice, not proof of a closed data license, a commercial AI evaluation transaction, or permission to disclose incident records.
Where might your own organization stand?
Take the private fit checkQuiz / Your next step
Are you assessing an incident corpus or a defined AI task?
A specific task and rubric can require different records and permissions than historical training data.
Your suggested next step
No fee for the initial conversation or introduction. We may be compensated by a buyer if an introduction becomes a partnership. No buyer, license, or payment is guaranteed. Review any proposed deal with your own legal and security advisers.