← INDEX

CASE

04 / 04

STATUS

COMPLETE

DURATION

3 WEEKS


ACCOUNTING · ENTERPRISE DATA CATALOG

Diagnosing a data catalog's labeling problem before a rebuild

A tagging system was quietly breaking search. Diagnosing the real cause first saved leadership from an expensive rebuild aimed at the wrong target.

01 OVERVIEW

The client's data catalog used tags to organize datasets, but labels came from two uncoordinated sources — system-generated and ad hoc employee entries — and search kept breaking down as a result. As the sole researcher on the Catalog team, I led a three-stage evaluative study to find out what was actually broken and turned the answer into a refined label list and set of recommendations that shaped the taxonomy rebuild that followed.

My role

  • Designed and led a 3-stage evaluative study to diagnose why the tagging system was breaking down
  • Recruited and moderated sessions with 9 participants across 3 service lines, 2 experience levels
  • Synthesized findings into a refined label list and recommendation log
  • Presented to leadership, shifting the plan from an open-ended rebuild to a scoped label refinement

02 CHALLENGE

How might we determine whether the catalog's search problems were a labeling issue or something bigger, before committing to a costly rebuild?

The scope of the problem was unclear. A rebuild aimed at the wrong root cause would be a significant, unrecoverable investment. Similarly, a recommendation made without evidence would carry little weight against competing priorities. And no one had yet asked the people using the catalog day to day, across experience levels and service lines, where the labels actually broke down for them in practice.

Goals

  • Quantify the extent of labeling ambiguity across user types and service lines
  • Understand how experts and novices actually think about labels
  • Produce a prioritized set of recommendations leadership could act on right away

03 APPROACH

A single test wouldn't show where the labels broke down, or why, but a sequence would.

n=9 participants · 5 expert (Sr. Associates, Managers) · 4 novice (Associate level) · 3 LoS — Tax, Assurance, Advisory

StageMethodResearch questions
01 Open Card Sort — users create their own categories, then sort cards into them Avg. # of categories created · how sorting shifts by experience/LoS · which labels were hardest to place, and why
02 Closed Card Sort — users sort a fixed set of new cards into Stage 1's categories, then write in and place any missing labels Avg. # of new labels written in · which categories absorbed them · how write-ins reshape the overall structure
03 Live Label Test — in a mid-fi prototype, users assign labels from the bank to sample datasets Which labels get chosen · avg. labels applied · requested-but-missing labels · reasoning behind label choices

04 FINDINGS

The catalog didn't need fewer labels. It needed sharper ones, and space for users to shape the rest.

01 · Users converged on 3–5 labels per dataset — never fewer, rarely more.

Below that, users couldn't find what they needed. Above it, the system read as noisy instead of precise.

02 · Experts and novices sorted differently, until labels got specific.

Novices grouped by broad, familiar categories; experts by domain workflow. That split disappeared once labels were tied to a single service line — both groups trusted those more. Structure needed to flex for experience; specificity didn't.

"I think the more granularity you get, the better the recall."

— Internal user, Platform Catalog

03 · Tax and Deals generated 40+ suggested labels, more than double any other line.

Everyone wanted to add their own tags, but requests concentrated where domain coverage was weakest. That became the prioritization signal for the refined label list.

"We don't have very useful tags — given what we do in our space, labels don't fully support what we do."

— Internal user, Platform Catalog

05 IMPACT

Leadership had been weighing a full rebuild of the search and catalog system. This research reframed the problem — labels, not infrastructure — shifting the roadmap to a scoped, evidence-backed fix.

  • Decision reframed — gave leadership the evidence to choose a scoped fix over a full rebuild
  • Partnership deepened — product and data teams aligned on the same diagnosis
  • Scope expanded — research extended into a second phase on tagging behavior

Deliverables

  • Label trend analysis by LoS
  • Refined label list, prioritized by gap size
  • Recommendation log grounded in IA heuristics

Next steps

  • Finalize categories from card-sort data
  • Validate the Tax list with the Catalogue Librarian
  • Audit ML for future label suggestion

06 REFLECTION

This project sharpened how I diagnose ambiguous "something feels broken" requests. What looked like a search infrastructure problem turned out to be an information architecture and content governance problem, and the only way to know that was to run structured evaluative research before recommending (or accepting) an expensive fix.

It was also early proof of the value of sequencing methods. No single test would have both surfaced the mental-model mismatch and validated a fix. That layered approach is something I still lead with today.

CASE_FILE: 04 / 04 WORD COUNT: ~680 LAST UPDATED: AUG 2026