Article URL: https://github.com/Chiaro-HQ/methodology Comments URL: https://news.ycombinator.com/item?id=49171140 Points: 3 # Comments: 0

This is the complete Chiaro methodology for SOC 2 readiness and audit: the control library we test against, the criteria each control maps to, the evidence we accept, and the rules the collection runs under. Readiness and the examination run on the same framework, so the bar a company prepares against is the bar the examination applies. It is published because the alternative to showing your work is asking people to take it on faith, and a compliance industry that ran on faith is why anyone is reading this. The attributes carry 498 worked examples of a judgment call between them, each recording a verdict an AI reached, the verdict that was correct, and why. They exist to calibrate judgment, and they are the part of this repository we would least like a competitor to have. They carry our experience, not our clients' information. Every published example carries source: "synthetic_taste_v1". Each one takes a judgment call we have actually had to make and writes it as a scenario for the purpose. That is deliberate, and it is the stricter of the two options: client work is confidential, and redaction removes a name without removing an identity, so a case lifted from fieldwork stays recognizable to anyone who knows the company. What is published here is the lesson: the distinction being drawn and the reasoning behind it. The companies, systems, numbers and people in the scenarios are not real and are not any client of ours. Which way do they push? 303 of them correct an AI that was too strict, and 195 correct one that was too lenient. We are publishing that ratio because anyone with the file can compute it, and because the honest reading is not flattering by default: an auditor whose examples mostly teach "that is fine actually" is exactly what the industry should be suspicious of after 2025. Our answer is that the two errors are not equally common. An engine reading evidence against written criteria over-flags far more often than it under-flags, because it has no way to see that a missing artifact is covered three ways elsewhere. Correcting that is most of the work. But the examples that push the other way are the ones that matter for an opinion, so we deliberately added to them rather than leaving the ratio where it fell, and we did not force it to 1:1, which would have been its own fiction. If you think a specific example softens something it should not, that is a concrete thing you can point at. Open an issue with the control and attribute id. That is the entire reason this file is public. By default we do not sample. The data a modern company runs on is produced by machines, so it can be verified at machine speed, all of it. The client's AI retrieves the complete population for each control (every change, every termination, every access review in the observation window). Completeness is corroborated: recorded retrieval always, and reconciliation against an independent second source wherever one exists. Then every item is tested, deterministically where the evidence is structured, by calibrated reading where it is prose, with every candidate deviation confirmed by the CPA before it becomes an exception. Sampling survives only where a population genuinely cannot be retrieved in full, and the report discloses, per control, which lane ran. When sampling does run, nobody picks: selection is seeded from a hash of the banked population itself, so neither side can steer or re-roll it.