Scribewave
Transcript Editor

Pseudonymization

Replace names and other identifying details with stable codes like P1, and check that nothing was missed.

Pseudonymization replaces identifying details in a transcript with short codes, and keeps the same code for the same person throughout the document:

"Myrthe says that Katy Perry looks good today, Jonah agrees." → "P1 says that Katy Perry looks good today, P2 agrees."

Whether the celebrity in that sentence also gets a code is a setting, not a guess — see Settings.

This is different from redaction. Redaction blacks out a name; pseudonymization keeps the transcript readable and analysable, because you can still tell who said what — you just can't tell who they are.

Where to find it

Open a project in the editor, then open the Pseudonymize panel (shield icon) in the right-hand sidebar, under Tools.

The workflow

1. Find

Click Find sensitive info. Scribewave reads the transcript part by part and shows its progress (Reading part 3 of 9…). Two things happen at once:

  • Patterns are matched exactly. Email addresses, phone numbers, ID numbers, bank accounts and card numbers are found by their shape, not by a model. IBANs and card numbers are checksum-validated, so these are effectively free of guesswork.
  • Speakers are seeded from the recording. Everyone the diarization identified as a speaker gets a code straight away, whether or not their name is ever said out loud.
  • The model finds the rest — people, organizations, places, and anything else you have switched on.

Run it again later (the button then reads Find more) after editing, or to pick up anything a first pass missed.

2. Review the code table

Every identity found becomes one row in the Code table: its code, the name it was recognised as, the category, how many mentions were found, and any other spellings it appeared in. Each row lets you:

  • Jump to the first mention, to see it in context.
  • Leave it uncoded — for a public figure, a company you don't need to hide, or a false positive. If it was already replaced, this puts the original text back.

Review at this level rather than mention by mention. A two-hour interview typically produces a few hundred replacements but only fifteen to forty identities, and it is the identities where the mistakes are: two different people merged into one code, or one person split across two.

Use the copy icon next to Code table to copy the whole table out, for your protocol or your data-management plan.

3. Replace

Click Replace n mentions with codes. Every occurrence of every known spelling is replaced across the whole document at once, and the speaker labels beside the paragraphs become codes too.

Replacements that Scribewave is not certain about are held back rather than applied, and collected under Needs your decision:

  • short forms and first names that are also ordinary words,
  • inflected forms (Myrthes, Jansens, Finnish or Polish case endings),
  • a name written in lowercase, where it is probably the common noun and not the name.

Each one shows its sentence so you can decide, or you can accept them all in one click. This is the deliberate trade: certain matches are silent, uncertain ones cost you a click and never a surprise.

4. Check for anything missed

From the ⋮ menu, choose Check for anything missed. This re-reads the transcript after replacement and asks a different question from the first pass — not "which words are names?" but "reading this as a whole, could a reader work out who this is?".

It reports two kinds of finding under Second-pass check:

  • missed — an identifier the first pass did not catch.
  • indirect — something that identifies a person without naming them: "the guy who lost his leg in the 2019 factory fire", "my sister who runs the bakery on Kerkstraat", a job title that only one person in a small organization holds. These are what make pseudonymized research data re-identifiable in practice, and span-level detection cannot see them.

For each finding you can show me (jump to it), give it a code, or mark it not a problem.

5. Verify

The panel keeps a running check on the document itself: for every identity in the code table, does any original spelling still appear anywhere? If one does, you get a warning naming it — "Myrthe" still in the transcript (3×) → P1 — and clicking it jumps you there. Speaker labels still carrying a real name are called out the same way.

This is a mechanical check of the text, not a model's opinion, so a clean result is a fact. What it cannot tell you is whether someone was never found in the first place — that is what step 4 is for.

Codes stay consistent

P1 in the first minute is P1 two hours later. Codes are assigned by Scribewave, never by the model, from a counter that runs across the whole transcript — so consistency is structural, not something the model is asked to remember. Codes are numbered in order of first appearance, which also makes the code table readable top to bottom against the recording.

Codes are scoped to one transcript. The same person in a different recording gets a code from that recording's own table.

Undoing it

Restore original names in the ⋮ menu puts the original text back everywhere, including the speaker labels. The original wording of each replaced mention is stored with the word itself, so Myrthe and mevrouw Jansen both come back as what they were, not as one canonical form.

Settings: who and what gets a code

⋮ → Settings. These apply to the whole organization, so only an admin can change them, and they apply to future runs — changing them does not rewrite work already done.

Who gets a code

  • Only participants — people who speak in the recording. Anyone merely talked about is left alone.
  • All private individuals — everyone named, whether they speak or not, but not public figures. This is the default.
  • Everyone named — including politicians, celebrities and other public figures.

What gets a code

Switch categories on or off and set the prefix each one uses. The prefix becomes the code: a prefix of P produces P1, P2, P3. Use R for respondent, PT for patient, or whatever your protocol asks for.

CategoryPrefixDefault
PeoplePOn
OrganizationsORGOn
Places and addressesLOCOn
Email addressesEOn
Phone numbersTOn
ID and passport numbersIDOn
Bank accounts and cardsACCOn
Dates tied to a personDOff
Job titlesROLEOff
AgesAGEOff
Other identifiersXOff

Prefixes are 1–8 characters, letters and digits only, and no two categories may share one — two categories both minting P1 would produce two different people with the same code.

Turn on Dates tied to a person and Job titles for health, legal and small-organization work, where a date of birth or a unique role identifies someone as surely as a name does.

Things worth knowing

Pseudonymization is not anonymization. Under GDPR Article 4(5), pseudonymized data is still personal data, because the mapping back to real names exists. Treat the code table as sensitive: a transcript exported together with its full key is not pseudonymized at all.

The audio is unchanged. The recording still says the name out loud. No text process can address that.

Original names survive elsewhere in the project — in edit history, in chat conversations about the transcript, and in the media file itself. If you need names gone rather than hidden, that is a separate request; talk to us.

Translations are pseudonymized separately. If you pseudonymize a translation and view it side by side with the original, the editor warns you that the reference column still shows real names.

Finding and checking use credits, because both read the transcript with a model. Replacing, restoring and editing the code table do not.

If two people work on the same transcript at once, whoever saves second is told "Changed in another session" and asked to reload the code table rather than overwriting the other person's work.

On this page