Abstract
<p>Asking crowds of stakeholders to generate and prioritize ideas has become commonplace, from citizen assemblies to corporate open-innovation challenges. Although the reliability of the results of such consultations is taken for granted, their consistency across independent groups has received little empirical attention, largely because such processes are typically one-off events that cannot easily be repeated. In this exploratory study, we introduce the concept of parallel-crowds reliability, the consistency of outcomes when the same collective intelligence process is applied to the same problem by different but demographically equivalent crowds, and we provide an initial empirical test of it. Two separate but equivalent groups of MBA students in the same course were asked to independently propose and prioritize ideas to improve their MBA program, with the same decision-maker curating and selecting winning ideas in both groups. We assess reliability at three successive stages using semantic similarity measures based on two independent embedding models. We find that the two groups independently proposed semantically similar ideas, with convergence concentrated among valuable ideas and divergence among poor ones. They also prioritized similar ideas similarly, enabling the decision maker to select similar sets of winning ideas. Given the single-domain design, we frame these results as an initial evidence base for future investigations.</p>