What to Audit Before Cleaning a Media Catalog
A pre-cleanup audit records the current catalogue state without editing it, creating the baseline needed to plan, reconcile, and reverse later changes.
In short: Freeze edits and record the current state: authorised sources and owners, collection groups, category counts, active filters, uncategorised items, empty categories, duplicate candidates, versions, language metadata, unavailable items, profile-specific state, and recovery options. Audit first; do not rename, merge, or delete while building the baseline.
An audit turns “the catalogue feels messy” into a bounded set of observable problems. It also provides the counts and examples needed to prove that later batches did not silently lose context.
Freeze the observation window
Choose a date and short period during which no one performs broad catalogue edits. Record:
- devices and app or browser versions used;
- account and profiles inspected;
- source labels;
- filters cleared or intentionally retained;
- audit start and end time;
- people conducting the audit.
A live source can still change. The time stamp makes that limitation visible.
Inventory sources and collection groups
Create one row per authorised source or major collection group. Include owner, access status, media types, broad scale, metadata strengths, known dependencies, and authorisation review state.
NARA inventory guidance describes an inventory as a descriptive listing of groups or systems, not necessarily every item. Its requirements apply to records, but the group-first approach prevents a personal audit from becoming an unplanned item-level cleanup.
Record categories and filters
For each visible category, note:
- exact label;
- source or system that defines it;
- displayed item count;
- clear household meaning;
- overlap with another category;
- last known use;
- whether active filters affect the count;
- whether unavailable or archived items may be hidden.
Do not remove an apparently empty category yet. Follow the empty category verification guide after the baseline.
Sample edge cases
Choose representative items:
- normal complete item;
- uncategorised item;
- possible duplicate;
- multi-version title;
- multilingual item;
- episode with incomplete metadata;
- unavailable item;
- archive candidate;
- item visible on one device but not another.
Record the source, version, category, language labels, availability, and profile context. The sample identifies rule needs; it does not prove the frequency of each issue.
Audit personal and shared state separately
Do not mix shared catalogue structure with:
- playback progress;
- viewing history;
- favourites;
- language preferences;
- profile-specific discovery.
Record only the minimum personal evidence needed for cleanup planning and protect privacy. A category cleanup should not clear another person's profile state.
Review recovery readiness
Document:
- supported catalogue export or snapshot;
- what the export contains;
- date of last successful creation;
- rollback or restoration test;
- source recovery owner;
- changes that cannot be reversed through the tool;
- approval needed before destructive work.
Use the reversible cleanup guide to turn this into a change-control process. Do not call an untested export a backup.
Original evidence: audit baseline sheet
| Area | Count/state | Observation method | Example | Risk/question |
|---|---|---|---|---|
| Sources | ||||
| Categories | ||||
| Uncategorised | ||||
| Empty categories | ||||
| Duplicate candidates | ||||
| Version groups | ||||
| Unavailable items | ||||
| Recovery |
Add audit date and filters. Feed the results into the complete cleanup plan and route uncertain items through the uncategorised triage method.
Distinguish observation from diagnosis
“Category has zero visible items with filters cleared” is an observation. “The category is safe to delete” is a decision requiring source and dependency checks. Keep those stages separate.
Likewise, matching titles are duplicate candidates, not proven duplicates. Missing language metadata does not prove a track is absent.
Common mistakes and limitations
- Editing while auditing.
- Recording counts with unknown filters.
- Listing every item before collection groups.
- Treating title matches as confirmed duplicates.
- Ignoring profile state.
- Calling an untested export a recovery path.
- Inferring cause from one screenshot.
- Auditing only the maintainer's preferred device.
Source data may change during or after the audit. Preserve time, device, and source context.
Frequently asked questions
Must every item be audited?
No. Inventory source and category groups, then sample edge cases. Item-level review is needed only for decisions such as duplicate merging or metadata correction.
Can cleanup begin while the audit continues?
Avoid overlapping broad work. Complete and approve a baseline for the current scope so later changes can be reconciled.
What is the most important audit output?
A dated baseline that another person can understand: what was included, how it was observed, which filters applied, and which issues remain questions.
Your next step
Explore Norva's catalog features
Sources
- National Archives: Records inventory introduction
- National Archives: Knowing your records
- Library of Congress: Inventory and custody
- Norva features