Over a billion years ago, a pair of massive black holes spiraled into each other. The resulting collision sent a gravitational wave propagating through the universe at the speed of light.
On September 14, 2015, that gravitational wave transient passed through the Laser Interferometer Gravitational-Wave Observatories (LIGO) at Hanford, Washington, and Livingston, Louisiana. Each LIGO observatory consists of an L-shaped tunnel, having 4-kilometer-long arms. At the intersection of the arms, an extremely powerful laser is split and sent down the length of each arm. Mirrors at the ends reflect the beam back to the origin. The gravitational wave that passed through each observatory literally stretched space, causing unimaginably tiny shifts in the relative distances between the origin point and the mirrors. The resulting differences in round-trip times produced minuscule but observable interference patterns when the beams recombined. Consistent, time-shifted interference patterns were observed at each observatory, providing the first direct observation of gravitational waves1.
In 1916, Albert Einstein predicted the existence of gravitational waves. In the 1960s, scientists attempted to detect them using room-temperature single-object resonant mass detectors. However, the single-object detector approach could not resolve a gravitational wave signal from the noise, even after improving the detector’s sensitivity by cooling it to near absolute zero.
The conceptual shift from seeking a central signal to reading the interference between paths may be what structured group deliberation needs.
Meanwhile, scientists began to contemplate an interferometer approach. As the name suggests, such instruments measure interference. That is, they seek the signal in differences. Ultimately, that’s what made detection possible. The conceptual shift from seeking a central signal to reading the interference between paths may be what structured group deliberation needs.
I’ve been thinking about that mindset shift since writing The Quiet Failure of Well-Run Meetings. To recap, the hidden-profile literature is unambiguous: groups systematically over-discuss the information and perspectives that members share and under-discuss those that only some hold. The structured deliberation methods that evolved over the last several decades were designed to counteract that neglect. They share a common architecture: independent generation before group discussion, anonymity to neutralize status effects, and aggregation before debate.
I argued that despite this shared architecture, most of these methods drift toward convergence — toward looking for signal. Their aggregation logic favors central tendency. That is, their deliverables are organized around themes, recommendations, or rankings. While the architecture captures divergence, the synthesis often doesn’t2.
While the architecture captures divergence, the synthesis often doesn’t.
The pattern I perceive is what I’ll call synthesis drift: methods whose input architecture faithfully captures divergence, but whose synthesis layer quietly filters it back out. The four methods below illustrate this drift to varying degrees. That’s not because they were poorly designed, but because any pressure to produce a deliverable that “looks like an answer” tends to pull synthesis toward central tendency. Viewing them through this lens makes the design choice behind my firm’s Reflect & Engage method easier to see when we get there.
Let’s take a look at what each of four established methods accomplishes. They are not failures. They are successful at what they were designed for. The challenge is understanding whether a particular tool is right for the job at hand.
Delphi emerged from RAND in the early 1960s3, designed to iterate toward consensus. Murray Turoff designed a variant intended to do the opposite, and the field largely ignored it.
Olaf Helmer and Norman Dalkey were trying to forecast technology with panels of geographically distributed experts whose status differences and personal animosities made productive group discussion nearly impossible. Their solution was elegant: anonymous written rounds, controlled feedback between rounds, and iteration toward stable estimates.
Delphi works. It has produced credible forecasts in technology assessment, public health policy, defense planning, and scenario development for sixty years. Its architecture neutralizes many of the social mechanisms that distort group judgment, and its multi-round structure addresses several causes of hidden-profile bias simultaneously.
But classical Delphi was designed for problems where convergence is the deliverable. It’s useful for forecasting when quantum computing will reach commercial viability, how many hospital beds the system will need by 2035, and what the probability of a pandemic outbreak is in the next decade. The multi-round iteration is not a side effect; it is the product. However, Bolger and Wright, in a careful 2011 critique, note that the convergence pressure can suppress valid minority views even when those views are based on information the majority lacks4. That's a known tradeoff. Sometimes, it’s worth accepting.
Within seven years of Dalkey and Helmer's published account, Turoff designed a variant intended to do the opposite, and it deserves more attention than it gets. His 1970 paper "The Design of a Policy Delphi" describes a method for generating the strongest possible opposing views.5 Panels are recruited explicitly for heterogeneity. Iteration is used to sharpen disagreements rather than smooth them. The policy Delphi was a divergence-preserving variant designed for policy questions, which usually address complexity not subject to unambiguous, expert-determined right answers.
The policy Delphi is the closest historical precedent for Reflect & Engage. Turoff developed the design direction more than 50 years before we showed up. Policy Delphi is also the path less traveled. When sponsors expect deliverables that look like answers, the divergence-preserving variant tends to lose ground to its consensus-seeking counterpart.
Polis emerged from deliberative-democracy work in the early 2010s. It was used, perhaps most famously, in the vTaiwan platform. It was designed to preserve opinion groups and respect minority dissent while simultaneously surfacing points of consensus between them. Participants vote agree, disagree, or pass on short statements; machine learning identifies clusters of opinion that share voting patterns; the visualization makes opinion groups legible without flattening them into a single position6.
Polis is closer to our own design intent than any other method I know of in regular use. Both make the multiplicity of views legible, but Polis treats this as a means toward finding group-informed consensus, whereas our method treats dissensus itself as the primary deliverable. Polis is designed for large-scale public deliberation, engaging up to hundreds of thousands of participants in policy questions in which the population itself is the unit of analysis. It would not be the right tool for a nine-person leadership team preparing for a decision on Friday. Due credit, though, for its innovative diversity-preserving design, scalable infrastructure, utilization of machine-learning techniques, and visualizations.
Outset represents a new generation of platforms that use generative AI to conduct adaptive interviews at scale and synthesize results7. Outset can run hundreds of interviews at once, in any language, probe for the why behind the answer, and quickly deliver themes, summaries, and highlight reels.
The technology is impressive. Researchers use it for concept testing, market segmentation, and user research. Within those contexts, theme-oriented synthesis is what’s needed. A product team trying to understand why a feature isn’t resonating with users wants to surface and understand recurring patterns.
By design, the tool seeks to listen for the dominant signal in the data, including recurring themes, persistent language, and emerging patterns. As a result, divergent views risk becoming outliers, then footnotes, then nothing at all. That’s appropriate when the question is “what do most users experience?” It’s the wrong approach when the question is “what does our group not yet know that it disagrees about?”
Delphi, Polis, and Outset are all scaled or technologically mediated approaches. The fourth method on this list belongs to a different lineage. It’s relational rather than computational, but it can also exhibit synthesis drift.
The standard practice for facilitator preparation is one-on-one confidential interviews. A skilled facilitator typically schedules thirty to sixty minutes with each participant, asks open-ended questions, and develops a conceptual map of the territory before convening the group. These interviews are effective. They are also logistically complex, time-consuming, and expensive.
When done well, this is good work. The relationship is the mechanism for depth. A skilled interviewer can follow an emergent thread, let a disclosure emerge from a pause, and ask the unscripted question that says, in effect, “that thing you just said matters.” This kind of relational thinking cannot be replicated by asynchronous formats8.
The trade-off is structural. The facilitator serves as the sole channel for pre-meeting intelligence. That is, they are simultaneously the instrument of data collection, the interpreter of meaning, and a participant in the subsequent group discussion.
More fundamentally, even when the facilitator is skilled, confidentiality is not the same as anonymity. The participant knows the facilitator knows who said what. This is a behavioral assurance, not a structural guarantee. It’s a promise, not a breakwater. Those who most need protection are often those with the most experience of unfulfilled promises9.
There is a second structural factor that can be a limitation. The interviewer’s prior hypotheses shape their questions. Information disclosed passes through a filter that the group never sees. Depending on the situation and the facilitator’s skill, this might be a benefit. Sometimes, it’s not. In either case, the filter is invisible to the group.
All of the preceding methods, including Reflect & Engage, share a foundational architecture. That is, they utilize independent input before group discussion, some form of anonymity or confidentiality, and synthesis. They differ in scale, modality, and iteration count, but those variations sit within the shared architecture. The key distinctions are at the level of purpose.
Reading the top row left to right traces a gradient from divergence-preservation to convergence by design. (The “Delphi” column refers to the dominant forecasting variant; Turoff’s policy Delphi would sit closer to Reflect & Engage and Polis.)
What’s interesting about these structured deliberation methods, taken together, is that their architecture tends to be more divergence-friendly than their synthesis. The independent-generation step in each is a faithful response to research on hidden profiles. Participants share their views without outside influences. Anonymity (or trusted confidentiality) does its work.
Then the synthesis takes those independent inputs and squeezes them through a process whose default behavior is to find the signal. Delphi iterates toward consensus. Outset surfaces themes. Interviews yield integration. Even Polis, which resists convergence at the visualization layer, does so by clustering.
In my experience, most skilled interviewers work hard to preserve minority perspectives, but doing so requires swimming against a structural current. The default, absent that deliberate effort, is integration.
Some of this tilt is methodological; it’s built into the aggregation logic. Some of it, I suspect, is sociological. Sponsors and clients of structured deliberation often expect deliverables that look like answers. Themes feel actionable. Recommendations feel decisive. A document that says “here’s where your group is operating from incompatible assumptions and doesn’t yet know it” feels less complete than a document that says “here are the three priorities everyone agrees on.” Intentionally or not, the synthesis layer can drift toward convergence.
What would it mean to read participants’ input the way LIGO reads the interference between its two arms?
It would mean treating the differences themselves as the deliverable. Not extracting a central tendency. Not filtering toward themes. Not iterating toward consensus. It would mean sensing where paths diverged and asking what produced the divergence.
The unit of analysis would shift. The question would no longer be “what does this group think?” but “where are members thinking differently?” The synthesis would not produce a forecast or a recommendation. It would produce a map of fault lines the group hadn’t yet noticed.
This is what Reflect & Engage is designed to do, and we’ve begun experimenting with it “in the wild” with select facilitators. The architecture is conventional, though it uses contemporary enabling technologies: anonymous asynchronous voice reflection, AI-supported synthesis, and a handoff to a facilitator. As noted above, the difference lies in the synthesis layer. The primary deliverable is not the group’s signal. It is the map of where the paths diverged: minority perspectives, unexamined assumptions, and tensions that risk being lost in a standard group process.
Good decisions require that the group develop a shared understanding of the situation. That is, the group orients based on complete situational awareness. Convergence is ultimately necessary. Premature convergence based on an incomplete map is the problem.
Our premise is that convergence is work for the facilitated meeting itself, performed by humans who can hold tension productively.
Our premise is that convergence is work for the facilitated meeting itself, performed by humans who can hold tension productively. Convergence is not the work of the pre-meeting structured deliberation method.
LIGO measures distance with extraordinary precision because the laser is coherent, the mirrors are stable, the vacuum is engineered, and the noise sources are well-understood. None of those conditions holds for group decision-making. People are not lasers. Reflection responses aren’t beams. The “interference pattern” produced by participant divergence is not a physical phenomenon; it is an analyst's interpretation, with all the bias and uncertainty that entails.
The analogy risks overpromising. Reflect & Engage does not produce nanometer-level measurements of group dynamics. Rather, we’re producing a structured account of where members appear to be operating from different assumptions, based on our (AI-assisted) reading of voice recordings and transcripts, and presenting it to a facilitator who must decide what to do with it.
And the analogy doesn’t address what happens after the reading. Group decision-making doesn’t end with the Field Report; it ends with a decision made by humans under conditions that can reintroduce all the social pressures the method was designed to remove. Whether the “interferometric reading” actually changes the conversation depends on the facilitator’s skill and the group’s willingness to engage with what’s been surfaced. Neither is guaranteed.
There is one further limit, separate from the analogy, that bears reexamining. LIGO doesn’t have to worry about what its lasers might say in the presence of a sympathetic listener. Group decision-making does. In an earlier post, I argued that removing the witness from the interview has a cost, and Reflect & Engage removes the witness by design. A skilled interviewer creates what Edgar Schein called a relational container: the conditions under which a participant can risk saying something they haven’t fully formed yet. Some of the most valuable input in any pre-meeting process comes from thoughts that would dissolve in a void but survive in the presence of someone who is visibly witnessing them. Asynchronous voice reflection trades the relational container for structural anonymity. It reduces the social cost of disclosure, but it also forfeits some of the conditions that make depth possible.
This complicates the interferometric framing. The “interference pattern” is only as informative as what participants offered to the system. If the container Reflect & Engage provides is shallower than the one a skilled interviewer creates, some of what would have been the most decision-relevant divergence may never make it into the data at all. The reading might be precise about what’s there and silent about what isn’t. I don’t think this argues against the design; single-channel interviewing has its own losses, often larger ones, but it does argue against overconfidence in the analogy. The detection question isn’t only “what can be read in the interference pattern?” but “what was rich enough to register in the first place?”
What I think the analogy gets right is the acknowledgment that it’s possible to build an architecture for independent input and still misread what it makes available by treating divergence as noise and chasing the central signal anyway.
What the preceding makes visible is that the architecture question and the synthesis question are separate. Most of the differentiating work happens at the synthesis layer, where the choice is rarely made explicit.
Most structured deliberation methods have the right architecture, including independent generation, some degree of anonymity, and aggregation before debate. Those features are nearly universal at this point. What varies is what the synthesis does with what the architecture captured. If the deliverable is themes, the synthesis is filtering for central tendency. If it’s a forecast, a recommendation, or a ranked list of priorities, it’s filtering for central tendency. If participants who held minority views were to read the synthesis and find that what they said was either absorbed into the consensus or footnoted as a “noted dissent,” the synthesis filtered for central tendency, regardless of what the input layer preserved.
That’s not necessarily wrong. When the question is “what does most of this group already share?” — a budget vote, a feature prioritization, a forecast where the deliverable is a number — convergence-oriented synthesis is exactly right, and tools like Delphi and Outset are well-suited to it. The question is whether the situation fits that kind of scenario.
Here’s where I’m landing for the time being: when the group is about to make a consequential decision and is converging toward a shared orientation, that’s the moment to ask whether what’s needed is a clearer central signal or a reading of the interference between paths that no one has yet compared.
Abbott, B. P., et al. (LIGO Scientific Collaboration and Virgo Collaboration). (2016). Observation of gravitational waves from a binary black hole merger. Physical Review Letters, 116(6), 061102. https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.116.061102. The 2017 Nobel Prize in Physics was awarded to Rainer Weiss, Barry Barish, and Kip Thorne for their contributions to LIGO and the observation of gravitational waves.
The research basis for why surfaced divergence improves complex decisions includes Scott Page’s work showing that cognitively diverse teams outperform homogenous groups on complex tasks. See Scott E. Page, The Diversity Bonus: How Great Teams Pay Off in the Knowledge Economy (Princeton University Press, 2017). https://press.princeton.edu/books/hardcover/9780691176888/the-diversity-bonus.
Dalkey, N., & Helmer, O. (1963). An experimental application of the Delphi method to the use of experts. Management Science, 9(3), 458–467. https://www.jstor.org/stable/2627117. The Delphi project began at RAND in the 1950s; the 1963 paper is the canonical published account.
Bolger, F., & Wright, G. (2011). Improving the Delphi process: Lessons from social psychological research. Technological Forecasting and Social Change, 78(9), 1500–1513. https://www.sciencedirect.com/science/article/abs/pii/S0040162511001430. Worth reading in full for its careful treatment of where Delphi’s convergence pressure helps and where it harms accuracy.
Turoff, M. (1970). The design of a policy Delphi. Technological Forecasting and Social Change, 2(2), 149–171. https://www.sciencedirect.com/science/article/abs/pii/0040162570901617. Turoff’s framing — that policy questions, by definition, lack expert-determined answers and therefore call for surfacing opposing views rather than averaging them — has aged remarkably well. The policy Delphi has had a small but devoted following. Petri Tapio’s Disaggregative Policy Delphi (2003) added cluster analysis to identify distinct viewpoint groups. Martin Steinert’s Dissensus Delphi (2009) reframed the aim as maximizing rather than minimizing variance. Hilary Linstone, Turoff’s longtime co-author, kept the tradition visible in Technological Forecasting and Social Change through decades of editorship. The bibliometric data nonetheless confirm that classical Delphi dominates practice by roughly an order of magnitude.
Small, C. T., Bjorkegren, M., Erkkilä, T., Shaw, L., & Megill, C. (2021). Polis: Scaling deliberation by mapping high dimensional opinion spaces. Recerca: Revista de Pensament i Anàlisi, 26(2). https://gwern.net/doc/sociology/2021-small.pdf.
Other examples of AI interviewing platforms include Listen Labs, Strella, Versive, and others. The category is genuinely new. Most of these platforms launched in the last few years. Their synthesis-towards-theme orientation may evolve in unanticipated directions.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.