Casebook

One identifier, kept on purpose. Every other one, absent.

A case may carry a date of birth — the single direct identifier the schema holds, from which the age at each test is computed. That makes the data protected health information, kept under a business-associate agreement rather than presented as de-identified. Everything else that would name a person has no column at all.

The one identifier, and why it is kept

When you enter a date of birth on a case, it is stored, and the age at each testing is computed from it rather than re-typed — one source of truth for the value that anchors every norm. It is a direct identifier, and storing it is a deliberate choice that makes this a system holding protected health information. It is kept in one structured column, encrypted at rest and in transit, written to a per-practice access log, and covered by a business-associate agreement.

What is still not in the schema

No name. No medical record number. No email, phone or address. Not in any table — and a test fails the build if a column with one of those names ever appears. The date of birth is the single, named exception to that test; every other identifier stays forbidden on every table.

A person is otherwise a pseudonymous code you choose, such as PT-0412, and the mapping from that code to a name lives entirely outside the system. The code is constrained to 32 characters and is checked against anything name-shaped, so a full name cannot be typed into it either.

The date of birth has exactly one home

It is filled only when you type it on the case. It is never scraped from an uploaded report — the parser still refuses to extract a date of birth even when the report prints it — and it is never allowed into a free-text field, where a date in prose is refused at the point of writing. The age at testing is shown in full at every age; it is a plain number computed from the birth date, not a normalized one.

An uploaded report is read once and dropped

A finished report carries the name, the date of birth and often the whole referral letter. The import endpoint extracts its text for that one request, parses it, returns proposals, and keeps none of it: no column, no cache, no copy on disk. Keeping the source text around “for auditing” would undo the whole design in one column, which is exactly how this usually goes wrong.

Free text is where the rest is kept out

Two free-text fields exist: a case summary at 2,000 characters and a score note at 300. Both are sized for “profile consistent with post-CRT slowing” rather than a case history that would drift into identifying detail.

Structured identifiers — an email, a phone or fax number, a record number, a web or IP address, a full calendar date — are refused at the schema layer. Name-shaped phrases are warned and never refused, and the reason is worth knowing: a detector strict enough to catch “Grace” also refuses “Rey Complex Figure”. The clinical vocabulary it checks against is built from the catalog, so it knows the instrument names a neuropsychologist actually types.

There is also a sweep — one request that reads every free-text field in your practice and reports what looks like an identifier, so the question “is there anything in here we should not have” is answerable rather than assumed.

The written argument, for counsel

What is stored, what is not, and what a business-associate agreement obligates is set out category by category in docs/data-minimization.md — written to be handed to counsel. It is not a legal conclusion and does not pretend to be one.

And leaving

The whole knowledge base exports as CSVs with a codebook describing every column, on any ordinary day. Closing a practice deletes the practice, its people, every case, every score and the audit log — which goes too, because kept after the cases have gone it is a list of code names and the times somebody read them.

What is stored, exactly — the page most worth arguing with.

Look at what it holds

More answers