Synthesis is the bottleneck. Most museums collect data; few turn it into actionable insight. I’m exploring tools that might help — this experiment tracks what works and what doesn’t.
Listening sessions produce hours of transcripts that rarely get the attention they deserve. You read once, maybe twice, highlighting what seems important. You develop a feel for what’s there. But you never quite shake the suspicion that you missed something — or that what you flagged reflects your priors more than the data.
I’ve been testing whether that process can be split differently.
Making sense of qualitative data is two operations pretending to be one. There’s extraction — identifying candidate moments in the transcript. And there’s judgment — deciding which candidates actually meet your threshold for inclusion. We conflate these because humans often do both simultaneously. But they’re separable. Extraction can be delegated; judgment cannot.
I ran an initial test using actual participant data from a Value Realization Collaborative (VRC) project — a MaP cohort in which museum teams conduct listening sessions to understand how their museums can support newcomers. The extraction step happens computationally. The judgment step — deciding whether each candidate belongs and ensuring we have the information needed to make decisions — remains with the museum professionals who are making choices about how their museum should operate.
Across 11 transcripts so far, the process has surfaced an average of 42 discrete moments per transcript. These are instances where someone described what they were thinking, feeling, or personal rules they follow when approaching a particular challenge or goal. (In this methodology, each moment is called a concept.) Counts ranged from 30 to 67, depending on session length and density. That’s the easy part to report. The harder question is whether it’s able to distinguish signal from noise.
Market research data and the kind of data we need aren’t the same. A market research firm might accept “I usually prefer X” as a comment that can contribute to an insight for the institution. We need something more specific: moments rooted in personal history that reveal how someone actually thought or felt while pursuing the goal we’re studying. “I always do X” is a general statement. “In that moment, I felt X” is interior cognition. The distinction matters because what people say they generally do and what actually moves them in specific moments often diverge.
The extraction mostly respected that boundary. Fewer generalizations slipped through than I’ve encountered in past attempts. The system was distinguishing between what someone did or felt in a specific moment and what they claimed to generally believe. I’ve since added a cross-checking step that flags potential concerns. Across all 11 transcripts, the initial extraction produced 454 concepts; the validator flagged 172 of them — about 38% — as needing review. That's still meaningful filtering: instead of reviewing 454 items cold, the team reviews 172 items with specific audit concerns already surfaced. More importantly, each flagged item reveals where the system is struggling, allowing us to refine the extraction further and reduce that percentage over time. I’m certain we can reduce the # of flagged items with some more iteration.
A comparison emerged, allowing me to test this impression. @Seán MacQueen at the Royal Alberta Museum had already manually reviewed several of their transcripts before we ran the extraction on the same material. Nothing he had found was missing. Every quote he’d highlighted, it highlighted. Every moment he had identified, it identified.
That’s good, but that’s completeness, not accuracy. Some of what was extracted needed correction — verb tense inconsistencies or occasionally losing the perspective-inhabiting quality that makes this data useful. The boundary between emotional reaction and inner thinking gets fuzzy in places (though that’s true for human reviewers too).
Ultimately, it’s easier to correct than construct. That is, a draft that’s 62% right and needs refinement is more valuable than a perfect extraction that takes so long you never get to the pattern-finding that actually generates insight. And, again, I’m certain we’re going to be able to improve the accuracy with more refinements in the coming weeks.
If extraction is delegable, what else in the process might be? Next week, we’re testing whether the same approach extends to affinity mapping — grouping concepts by what participants were trying to accomplish when those thoughts and feelings arose. I expect this will actually be much simpler than identifying concepts that meet our criteria.
If you’re doing qualitative work and wondering whether this kind of process split could help: the question isn’t whether the extraction is perfect. It’s whether the draft is good enough to judge. That’s a different bar, and one worth testing against your own quality standards.