| N tasks | Main effect, K=2 | Main effect, K=3 | Conditional AMCE: attribute × moderator (N2, binary) | Attribute × attribute interaction (reference) |
|---|---|---|---|---|
| 3 | 1454 | 2181 | 8721 | 13082 |
| 4 | 1091 | 1636 | 6541 | 9812 |
| 5 | 873 | 1309 | 5233 | 7849 |
| 6 | 727 | 1091 | 4361 | 6541 |
3 Questionnaire design A: Distributive Conjoint Survey Experiment
3.1 Overview
This document specifies the design and implementation of the distributive conjoint module included in the survey. In this module, respondents act as members of a university committee and divide a fixed scholarship fund between two applicants whose attributes are experimentally randomized, so that allocating more to one applicant necessarily means allocating less to the other. The design adapts the Distributional Survey Experiment of Gilgen (2022) to a conjoint format (Hainmueller et al., 2014), combining the tabular presentation and complete, independent randomization of a conjoint experiment with the fixed-sum allocation task characteristic of distributive designs. The applicant attributes operationalize deservingness criteria drawn from the welfare deservingness literature, specifically the CARIN and recast NICER frameworks (Knotz et al., 2022; Meuleman et al., 2020; Oorschot, 2000); the theoretical rationale for each attribute, the derivation of the hypotheses, and the analysis plan are documented separately in the Pre Analysis Plan. The remainder of this document describes the survey structure, the attributes and their levels, the randomization procedure, and the technical implementation of the module.
3.2 Experimental design
3.2.1 Design type
The study uses a paired-profile distributive conjoint, a design that fuses the randomization logic of conjoint survey experiments with the allocative logic of distributive justice experiments. As in a standard conjoint, each task presents two profiles whose attributes are independently randomized, and the analysis recovers the effect of each attribute and level on the outcome. What departs from the standard conjoint is the response: rather than choosing one profile over the other, the respondent distributes a fixed resource between them, allocating a percentage of the total to each. The design therefore functions as an allocator—it asks how a scarce good should be divided between two independent claimants—and reveals the weight respondents place on each attribute through the share they are willing to grant.
Formally, each task displays two applicants with randomized profiles, and the respondent divides a fixed scholarship fund between them, so that the share allocated to one and the share allocated to the other sum to 100% by construction. The outcome is the percentage assigned to each profile. This fixed-sum response is what makes the design distributive rather than evaluative: because the resource is scarce and interdependent, any share granted to one applicant is necessarily withheld from the other, forcing respondents to reveal how they prioritize competing criteria when the trade-off is concrete.
This distributive logic follows the Distributional Survey Experiment of Gilgen (2022), but the present design differs from hers in two consequential respects. First, it retains the full randomization of a conjoint rather than constructing a D-efficient design: attribute levels are drawn independently and uniformly at the profile level, which simplifies both implementation and the identification of AMCEs (see Causal identification). Second, it presents two profiles per task rather than three. Using two profiles reduces the cognitive load of a task that is already more demanding than a binary choice, and it keeps both the interface and the statistical model simple, since the interdependence of the outcome becomes trivial: the second applicant’s share is the complement of the first (B = 100 − A). The result is a design that borrows the causal-inferential machinery of the conjoint tradition (Hainmueller et al., 2014) and the distributive, scarcity-based task of Gilgen’s framework, without inheriting the design-efficiency complications of the latter.
3.2.2 Attributes and levels
These are the criteria, attributes, and levels used to build each applicant profile. Levels are numbered so that higher numbers correspond to the more deserving level, with level 1 as the reference category. Identity is the exception: its numbering places the reference category (Chilean-born) first and is not a deservingness ordering.
| Attribute | Criterion | Profile label | Levels |
|---|---|---|---|
| Need | Need (CARIN/NICER) | Makes ends meet with (Su hogar llega a fin de mes con:) | 1 Comfort (Holgura) 2 Hardship (Dificultad) |
| Identity | Identity (CARIN/NICER) | Country of birth (País de nacimiento:) | 1 Chile 2 Peru (Perú) 3 Venezuela |
| Control | Control (CARIN/NICER) | Needs the scholarship because (Requiere la beca porque:) | 1 Did not apply to other scholarships in time (No alcanzó a postular a tiempo a otras becas) 2 Applied to other scholarships but received no funding (Postuló a otras becas pero no obtuvo financiamiento) |
| Effort | Effort (NICER) | Studies (Estudia:) | 1 Less than their peers (Menos que sus compañeros) 2 The same as their peers (Igual que sus compañeros) 3 More than their peers (Más que sus compañeros) |
| Reciprocity | Reciprocity (CARIN/NICER) | Outside their studies (Fuera de sus estudios:) | 1 Has not done volunteer work (No ha hecho voluntariado) 2 Has done volunteer work (Ha hecho voluntariado) |
| Attitude | Attitude (CARIN) | Sees the scholarship as (Ve la beca como:) | 1 Something they deserve (Algo que se merece) 2 Help they are grateful for (Una ayuda que agradece) |
| Sex | Ascriptive (signaled by name) | First name (no explicit row) | 1 Male name 2 Female name |
3.2.3 Scenario
Respondents are asked to place themselves in the role of a member of a university committee responsible for allocating a first-year higher education scholarship. They are told that the two applicants shown in each task were admitted to the same program at the same institution, where students must pay to study, that the information about each applicant was provided and verified by their secondary school, and that the scholarship fund available to distribute is CLP 2,000,000 for the first year. This framing establishes a context of genuine scarcity—the fund cannot fully cover both applicants—and holds the admission bar and institution constant across applicants, so that the allocation decision turns on the experimentally manipulated attributes rather than on differences in eligibility or institutional prestige.
3.2.4 Task structure
Each respondent \((i)\) completes six allocation tasks \((t)\). In every task, two applicant profiles (each labeled by a first name and displayed side by side in a table) are presented together with the fixed scholarship fund, and the respondent distributes the fund between them according to what they consider fair, with no correct or incorrect answer. Presenting six tasks per respondent rather than a single one increases within-respondent statistical efficiency by yielding multiple allocation decisions per person. This gain comes at the cost of non-independence among the tasks and profiles evaluated by the same respondent, which the analysis addresses explicitly (see Analysis Plan).
Six tasks sits well within the range the evidence supports: Bansak et al. (2021a) find no detectable degradation in response quality up to thirty tasks in online panels. It also matches practice in distributive designs, since Gilgen (2022) administers four allocation tasks with three profiles each—twelve profile-level observations per respondent, the same number generated here. Because the distributive task is more demanding than a binary choice, we treat thirty as an upper bound rather than a target.
3.2.5 Profiles and attributes
Each applicant profile \((j)\) is defined by seven experimentally manipulated dimensions. Six correspond to deservingness criteria from the NICER and CARIN frameworks (Knotz et al., 2022; Meuleman et al., 2020) and are displayed as explicit rows of the profile table: need, control, effort, reciprocity, attitude, and identity. The seventh, applicant sex, is signaled implicitly through a gendered first name rather than as an explicit row, so that the profile reads as a natural description of a person rather than a mechanical list of traits; names are drawn from pools matched on familiarity and perceived social class to isolate the sex signal from other connotations. The order in which the six explicit attributes appear is randomized once per respondent and held constant across their six tasks, so that any effect of attribute position is balanced across the sample while within-respondent presentation remains stable. The full set of levels and reference categories is specified in the Variables section.
Figure 3.1 shows the module as presented to respondents.
3.2.6 Causal identification and assumptions
The distributive conjoint identifies the average marginal component effect (AMCE) of each attribute on the allocated share under the design-based assumptions formalized by Hainmueller et al. (2014). Because attribute levels are assigned by the researcher rather than observed, identification rests on the randomization itself rather than on selection-on-observables, provided four conditions hold. We adopt their notation: let \(i\) index respondents, \(t \in \{1,\dots,6\}\) tasks, and \(j \in \{1,2\}\) the profile position within a task. Let \(T_{ijt}\) denote the vector of randomized attributes of profile \(j\), with \(l\)-th component \(T_{ijt}^{(l)}\) taking values in the level set \(\mathcal{T}_l\), and let \(Y_{ijt}(t)\) be the potential share the respondent would allocate to that profile if its attribute vector were fixed at \(t\).
Stability and no carryover effects. The potential outcome of a profile depends only on that profile’s own attributes, not on the profiles seen in the same or previous tasks. Formally, for any two tasks \(t\) and \(t'\) and profiles \(j\), \(j'\),
\[Y_{ijt}(t) = Y_{ij't'}(t) \quad \text{whenever the profile's attribute vector is the same,}\]
so that a profile with attribute vector \(t\) yields the same potential allocation regardless of when or against which competitor it appears. This licenses pooling the six tasks per respondent into a single estimand rather than treating each as a separate experiment.
No profile-order effects. The potential outcome does not depend on the position (left or right, here altID 1 or 2) that the profile occupies. For any two positions \(j\) and \(j'\),
\[Y_{ijt}(t) = Y_{ij't}(t),\]
which we enforce by randomizing which profile occupies each position, making position orthogonal to attribute content.
Randomization (completely independent). The assigned attribute vector is statistically independent of the potential outcomes and of the competing profile’s attributes. Writing the full profile of both alternatives in a task as \(T_{it}\),
\[Y_{ijt}(t) \perp\!\!\!\perp T_{it} \quad \text{for all } t,\]
and the components are assigned independently of one another with fixed marginal probabilities,
\[P\!\left(T_{ijt}^{(l)} = t_l\right) = p_l(t_l), \qquad \sum_{t_l \in \mathcal{T}_l} p_l(t_l) = 1,\]
with each attribute drawn uniformly, \(p_l(t_l) = 1/|\mathcal{T}_l|\), independently across the \(L\) attributes. This holds by construction, since levels are drawn independently and uniformly at the profile level.
Positivity (common support). Every attribute level, and every profile-level combination that the design admits, has strictly positive probability of assignment:
\[P\!\left(T_{ijt} = t\right) > 0 \quad \text{for all admissible } t \in \mathcal{T}_1 \times \cdots \times \mathcal{T}_L.\]
The only combinations excluded by the design are pairs of profiles within a task that are identical or that differ on a single attribute (see Randomization). Because this restriction operates on the joint distribution of the two profiles in a task and never removes any individual level from any attribute’s support, positivity at the attribute level is preserved.
Under these four conditions the AMCE of moving attribute \(l\) from a reference level \(t_l^{0}\) to an alternative level \(t_l^{1}\) is nonparametrically identified as
\[\text{AMCE}_l = \mathbb{E}\!\left[\,Y_{ijt}\!\left(t_l^{1}, T_{ijt}^{(-l)}\right) - Y_{ijt}\!\left(t_l^{0}, T_{ijt}^{(-l)}\right)\right],\]
where \(T_{ijt}^{(-l)}\) denotes the remaining attributes and the expectation is taken over their design distribution. The estimand is thus the average change in allocated share associated with the level contrast, averaging over all other attributes as randomized.
3.3 Variables
3.3.1 Attributes
The full set of attributes, their levels, and the numeric coding used in estimation is given in Section 3.2.2. Reference categories are set to the least deserving or baseline level of each attribute (level 1), so that estimated AMCEs read as the change in allocated share associated with moving toward the more deserving level. These reference categories are fixed in advance and applied identically in the analysis code.
3.3.2 Outcome
The outcome is the share of the scholarship fund allocated to each applicant. Respondents register their decision using a single continuous slider displayed beneath the profile table, initialized at an even split. The slider spans the full fund in Chilean pesos (CLP 0 to CLP 2,000,000), and as it moves, the percentage and peso amount for each applicant update in real time. The recorded value is the amount allocated to the right-hand applicant; the left-hand applicant’s amount is its complement, so the two always sum to the fixed total. For analysis this amount is converted to a share of the fund, bounded between 0 and 100, with the two shares within a task summing to 100 by construction. This fixed-sum structure is the defining feature of the distributive design: it makes the allocation genuinely zero-sum, so that any share granted to one applicant is withheld from the other. Modeling the outcome as a share rather than a raw amount respects this structure and follows the strategy recommended by Gilgen (2022) for distributional survey experiments.
3.3.3 Controls and moderators
Because attributes are randomized at the profile level, no covariates are required to identify the AMCEs: respondent-level characteristics are not confounders and enter only to estimate heterogeneity in the effects, never as adjustments to the main estimand (Hainmueller et al., 2014). The confirmatory moderation hypotheses (H8–H10) rest on two sets of respondent-level measures. The first is meritocratic orientation, measured following Castillo et al. (2023), which distinguishes four components: meritocratic perception (the belief that effort and talent are in fact rewarded), meritocratic preference (the normative endorsement that they should be rewarded), perception of privilege (the recognition that non-meritocratic factors shape outcomes), and acceptance of privilege (the normative endorsement of privilege-based advantage). The two normative components serve as the confirmatory moderators, since they yield directional predictions; the two descriptive-perception components are reserved for the exploratory hypotheses (H11–H12). Following Leeper et al. (2020), subgroup and moderation analyses rely on marginal means rather than differences in conditional AMCEs, which depend on the reference category and can mislead when compared across subgroups; the formal treatment is given in the Analysis Plan.
3.4 Randomization
3.4.1 Strategy
The experiment uses complete and independent randomization of attribute levels, the reference case in Hainmueller et al. (2014) and the default for most applied conjoint designs (Bansak et al., 2021b). For each profile \((j)\), every attribute is drawn independently of the others and of the competing profile, with uniform probability within its attribute (\(1/2\) for the four two-level attributes, \(1/3\) for Effort and Identity). No attribute is conditioned on another and no level is reweighted, so the joint distribution of profiles is the product of the attribute marginals (Hainmueller et al., 2014). Randomization occurs at the profile level and is performed anew for each respondent and task, rather than as a fixed block assigned in advance.
Under this scheme the assignment of each attribute is by construction independent of the potential outcomes, which is the condition that identifies the AMCE (see Causal identification), and no attribute can confound another. Crucially, identification does not require that every possible profile, or every possible pair, appear in the data: the AMCE is a marginal quantity, defined as the average effect of moving one attribute from its reference level to another while averaging over the distribution of the rest (Hainmueller et al., 2014). It is therefore identified from the marginal variation in each attribute, not from joint coverage of the factorial space. At the realized scale this coverage is overwhelming: the 54,000 profiles observed draw on a space of only 288 unique combinations, so each possible profile appears roughly 188 times, and each level of every attribute is observed between 18,000 and 27,000 times across widely varying combinations of the other attributes. That density, not the enumeration of the factorial or pair space, is what a marginal effect requires.
A direct consequence of independent randomization is orthogonality: across the sample, the level of any one attribute is statistically independent of the level of every other, so the estimated effect of one attribute cannot be biased by the distribution of another. Exact zero correlation holds only in expectation, so it is worth stating what “independent” means at a finite sample. With 54,000 profiles the sampling standard error of a correlation between any two attribute indicators is approximately \(1/\sqrt{54{,}000} \approx 0.004\), so inter-attribute correlations should fall within roughly \(\pm 0.013\) of zero (three standard errors) purely by sampling variation. A simulation of the design under the realized sample size confirms this directly: across the full attribute set the largest observed within-profile correlation is \(0.011\), within the expected band and stable across repeated seeds (see ?sec-random-check). Departures of this magnitude are sampling noise, not structural dependence, and bias no AMCE.
This is why a D-efficient design is not needed here. Such designs optimize the attribute covariance structure to minimize coefficient variance, and are valuable mainly when the factorial space is large relative to the sample, when specific interactions are the target, or when a forced-choice model is fit on few tasks—none of which binds in this design (Auspurg & Hinz, 2016; Gilgen, 2022). The space is small (288 profiles), the estimands are marginal main effects, and complete randomization already guarantees in expectation the orthogonality a D-efficient design engineers, delivering unbiased AMCEs with near-optimal efficiency for main effects at this size (Bansak et al., 2021b), without the added complexity.
3.4.2 Restrictions
The design imposes a single restriction, applied to the pair of profiles within a task rather than to individual attributes: the two profiles must differ on at least two of the six CARIN/NICER attributes (need, identity, control, effort, reciprocity, attitude). When a pair is generated, the six substantive attributes are drawn independently and the pair is retained only if it differs on two or more of them, otherwise it is redrawn. Sex is assigned afterward through the applicant’s first name and does not enter the check, so the two profiles may share the same sex or not. The purpose is cognitive: a pair that is identical or differs on only one substantive attribute presents two near-indistinguishable applicants, yielding a task that carries little information about trade-offs.
The restriction operates only on the joint distribution of the two profiles and leaves each attribute’s marginal distribution unchanged: every level still appears with its original uniform probability and remains independent of the other attributes within a profile. Because AMCE identification depends on this marginal randomization and not on the joint distribution of pairs, identification and interpretation are left intact; the restriction rules out only the least informative comparisons. Empirical attribute balance and orthogonality are verified as design checks (see ?sec-random-check).
3.4.3 Combinatorics
The table below reports the size of the design space and the coverage the realized sample achieves. The distinction that matters is between the profile space, over which the AMCE is defined, and the pair space, which is a by-product of showing two profiles per task and plays no role in identification. The admissible-pair count reflects the restriction explained above.
| Quantity | Value | Derivation |
|---|---|---|
| Unique profiles | 288 | 2·3·2·3·2·2·2 (incl. sex) |
| Ordered pairs | 82,656 | 288 · 287 |
| Unordered pairs | 41,328 | 288 · 287 / 2 |
| Pairs excluded by the restriction | 2,448 | 288 · 17 / 2 (differ on 0 or 1 of the six CARIN/NICER attributes, any sex) |
| Admissible pairs | 38,880 | 41,328 − 2,448 |
| Profiles observed (N = 4,500, 6 tasks) | 54,000 | 4,500 · 6 · 2 |
| Mean appearances per profile | ≈ 188 | 54,000 / 288 |
| Observations per level (2-level attr.) | 27,000 | 54,000 / 2 |
| Observations per level (3-level attr.) | 18,000 | 54,000 / 3 |
| Tasks observed | 27,000 | 4,500 · 6 |
| Distinct admissible pairs seen | ≈ 19,470 (50%) | expected coverage of 38,880 over 27,000 draws |
Two facts follow. First, the profile space is covered densely: each of the 288 profiles appears roughly 188 times, and every attribute level is observed between 18,000 and 27,000 times, which is what identifies the AMCE. Second, the pair space is covered only partially—about half of the admissible pairs are ever seen—but this is immaterial, since the AMCE averages over the marginal distribution of each attribute, not over the joint distribution of pairs (see Causal identification). Full enumeration of the pair space is neither necessary nor feasible.
3.5 Analysis Plan
3.5.1 Estimands
The primary estimand is the average marginal component effect (AMCE) of each attribute on the share of the fund allocated to a profile, in percentage points. For attribute \(l\), moving from reference level \(t_l^0\) to level \(t_l^1\),
\[\text{AMCE}_l = \mathbb{E}\!\left[Y_{ijt}\!\left(t_l^1, T_{ijt}^{(-l)}\right) - Y_{ijt}\!\left(t_l^0, T_{ijt}^{(-l)}\right)\right],\]
the average change in allocated share associated with the level contrast, averaging over the design distribution of the remaining attributes (Hainmueller et al., 2014). Since the outcome is a share of a fixed fund rather than a choice probability, the AMCE reads directly as the percentage points of the scholarship a level shifts toward or away from a profile. Alongside AMCEs we report marginal means—the average share at each attribute level—which describe absolute support without a baseline category and are the basis for all subgroup comparisons (Leeper et al., 2020).
3.5.2 Models
3.5.2.1 Main effects
The confirmatory estimator for the main effects (H1–H7) is a linear model of the allocated share on the seven attributes, with respondent fixed effects and standard errors clustered by respondent:
\[\text{share}_{ijt} = \alpha_i + \sum_{l=1}^{7} \beta_l\, T^{(l)}_{ijt} + \varepsilon_{ijt},\]
where each attribute enters as a set of dummies against its pre-registered reference category, so that each \(\beta_l\) is the corresponding AMCE. The fixed-sum structure implies that every respondent’s mean allocated share is exactly 50 by construction, so \(\alpha_i\) does not absorb between-respondent differences in baseline generosity—there are none to absorb. Its role is to remove the purely chance variation in the attribute composition each respondent happened to face, while the substantive work of accounting for non-independence is done by clustering, which is robust to the arbitrary within-respondent covariance the fixed-sum constraint induces, including the exact negative correlation between the two shares within a task. Modeling the share rather than the raw peso amount respects that structure and follows the strategy adopted by Gilgen (2022) for distributional survey experiments (Hainmueller et al., 2014).
3.5.2.2 Conditional effects
The confirmatory moderation hypotheses (H8–H10) target the conditional AMCE (Hainmueller et al., 2014): the same attribute contrast evaluated within the subpopulation defined by a respondent-level moderator \(M_i\). Formally, for attribute \(l\) moving from reference level \(t_l^{0}\) to level \(t_l^{1}\),
\[\text{AMCE}_l(m) = \mathbb{E}\!\left[Y_{ijt}\!\left(t_l^{1}, T_{ijt}^{(-l)}\right) - Y_{ijt}\!\left(t_l^{0}, T_{ijt}^{(-l)}\right) \,\Big|\, M_i = m\right].\]
Identification requires only that \(M_i\) be a pre-treatment respondent characteristic, which holds by design. Because randomization is performed independently of respondent identity, the design distribution of the remaining attributes \(T_{ijt}^{(-l)}\) is identical for every value of \(M_i\): unlike observational subgroup analysis, the distribution over which the effect is averaged does not differ across groups. What differs is the effective information available per conditional effect, which is the basis of the sample-size requirement for heterogeneity (see Power analysis).
Each moderation hypothesis is estimated with the same respondent fixed-effects specification used for the main effects, adding the interaction between the moderated attribute \(l^{*}\) and the standardized moderator:
\[\text{share}_{ijt} = \alpha_i + \sum_{l=1}^{7} \beta_l\, T^{(l)}_{ijt} + \delta\left(T^{(l^{*})}_{ijt} \times M_i\right) + \varepsilon_{ijt},\]
with standard errors clustered by respondent. The conditional AMCE of attribute \(l^{*}\) at moderator value \(m\) is then \(\text{AMCE}_{l^{*}}(m) = \beta_{l^{*}} + \delta m\), so that \(\delta = \partial\,\text{AMCE}_{l^{*}} / \partial M\) is the parameter of interest: because \(M_i\) is standardized to mean zero and unit variance, \(\beta_{l^{*}}\) is the AMCE at the sample mean of the moderator and \(\delta\) is the change in that AMCE per one-standard-deviation increase.
Although \(M_i\) is constant within respondent and its main effect is absorbed by \(\alpha_i\), the interaction remains identified. Writing \(\bar{w}_i\) for the respondent-specific mean of any variable \(w\), the within transformation applied by the fixed-effects estimator gives \(M_i - \bar{M}_i = 0\) for the moderator alone, but
\[T^{(l^{*})}_{ijt} M_i - \overline{T^{(l^{*})}_i M_i} = M_i\left(T^{(l^{*})}_{ijt} - \bar{T}^{(l^{*})}_i\right),\]
which is non-degenerate whenever the attribute varies within respondent—as it does by construction, since each respondent evaluates twelve independently randomized profiles. Estimating cross-level interactions within a fixed-effects specification is the strategy Gilgen (2022) adopts for heterogeneity in distributional survey experiments, interacting vignette-level attributes with respondent-level characteristics while retaining respondent fixed effects.
The absorbed main effect of the moderator is of no substantive interest here. Because the two shares within a task sum to the fixed total, every respondent’s mean allocated share is exactly 50, that is \((2T)^{-1}\sum_{t}\sum_{j}\text{share}_{ijt} = 50\) for all \(i\), so no respondent characteristic can predict it and level-2 main effects are mechanically null. For the same reason, a random-intercept specification would model the one quantity that carries no between-respondent variation and is therefore not used: respondent heterogeneity in this design resides in how strongly individuals weight each attribute, not in a baseline level of generosity.
One model is fitted per moderation hypothesis, with a single interaction term, so that each \(\delta\) is interpretable and the moderators do not compete for the same variance. The exploratory hypotheses (H11–H12) use the identical specification with the two descriptive-perception components as moderators. Following Leeper et al. (2020), any description of differences across subgroups is reported as differences in marginal means computed directly from the data, never by comparing conditional AMCEs across groups, since the latter depend on the choice of reference category.
3.5.3 Inference criteria
Hypotheses are evaluated two-sided at \(\alpha = 0.05\), with 95% confidence intervals reported for all AMCEs and marginal means. For the two three-level attributes, Effort and Identity, the corresponding hypotheses (H2 and H6 for main effects, H8 for moderation) are assessed with a joint Wald test of the two non-reference coefficients rather than coefficient by coefficient. The joint test is invariant to the choice of reference category, whereas the individual coefficients—and, in the moderation case, the individual interaction terms—are not (Leeper et al., 2020). Level-specific AMCEs are still reported for interpretation, with the reference category stated explicitly.
Each hypothesis therefore contributes exactly one \(p\)-value, and across the confirmatory family (H1–H10) we control the family-wise error rate with the Holm–Bonferroni procedure over those ten tests, reporting both unadjusted and adjusted \(p\)-values. Directional hypotheses (H1–H5, H7) are assessed against their pre-registered sign; identity (H6) is non-directional. Descriptions of differences across subgroups are reported as differences in marginal means rather than as comparisons of conditional AMCEs (Leeper et al., 2020). Exploratory analyses (H11–H12, the latter applying the same joint-test rule to Effort, and any socio-demographic contrasts) carry no confirmatory weight and are labeled as such.
3.5.4 Missing data
The slider carries a default even-split value and each task must be completed to advance, so item non-response on the outcome is not expected. Respondents failing the pre-registered quality checks (see Quality control) are excluded from the confirmatory analysis, with all results reported both including and excluding flagged respondents. Partial responses from respondents who abandon before completing all six tasks are retained for completed tasks, since the profile-level estimator does not require balanced task counts. No imputation is performed on the outcome.
3.6 Power analysis
This section asks a single question—is a sample of 1,500 respondents large enough to test the hypotheses?—and answers it twice, on purpose. We first apply a standard textbook formula for conjoint power, and then check it against a simulation of our actual design. The two are reported together because they play complementary roles. The formula is conservative by construction: it was derived for a different kind of outcome than ours and, as explained below, systematically asks for more respondents than we truly need. The simulation removes that conservatism and gives the realistic answer. The short version of the answer is that the design is comfortably powered for all main effects under either method, and that its ability to detect the moderation effects depends on how large those effects truly are—a point we make precise below.
3.6.1 Why two methods
The standard power formula for conjoint experiments (Schuessler & Freitag, 2020) was built for the usual conjoint task, where the respondent picks one of two profiles and the outcome is binary: chosen or not. A binary outcome has a fixed, and maximal, amount of variability, and the formula bakes that maximal variability into every calculation. Our outcome is not binary. Respondents move a slider to divide a fixed sum, producing a continuous share from 0 to 100. In practice these shares cluster around the middle rather than spreading across the whole range, so their variability is much lower than the binary worst case the formula assumes. Lower outcome variability means more statistical power for the same sample. The formula, blind to this, keeps charging us the worst-case variability and therefore overstates how many respondents we need. This is why we treat the formula as a conservative upper bound and use the simulation, which reproduces the real continuous outcome, to find the lower bound.
3.6.2 Analytic benchmark: the conservative upper bound
Under the binary-outcome formula produced by Schuessler & Freitag (2020), the number of respondents required to detect an effect of a given size follows
\[N \approx \frac{K}{2} \cdot \frac{(z_{1-\alpha/2} + z_{\kappa})^2}{\delta_1^2},\]
where \(K\) is the number of levels of the attribute, \(\delta_1\) the target effect, here 3 pp, the lower range of effects in the deservingness literature (Gilgen, 2022; Knotz et al., 2022), and the middle term equals 7.84 at the conventional \(\alpha = 0.05\) and 80% power. Two features carry through the whole section. Three-level attributes (Effort, Identity) are the demanding ones, because each of their contrasts uses only two-thirds of the profiles and so needs more respondents than a two-level attribute. And interactions—including the conditional AMCEs behind the moderation hypotheses—are far more demanding than main effects: detecting how an effect changes across groups requires roughly four times the sample of the effect itself, because the estimate now depends on the joint variation of two factors rather than one (Schuessler & Freitag, 2020, p. 9).
The table gives the resulting minimum sample, by number of tasks per respondent and by type of estimand.
Read the table as an upper bound, not as a verdict. Three things stand out. More tasks per respondent lower the required sample proportionally, since power depends on the total number of profile evaluations (respondents × tasks × 2). The jump from main effects to interactions is large: at six tasks, a three-level main effect needs about 1,090 respondents, while the same attribute interacted with a binary moderator needs about 4,360. And the attribute-by-attribute column is shown only for reference—those interactions are not hypotheses of this study, whose moderation hypotheses all involve a respondent characteristic, not two attributes. Taken at face value, this benchmark says that 1,500 respondents are more than enough for every main effect but fall short for the moderation effects. That verdict, however, rests entirely on the binary worst-case assumption. The next subsection removes it.
3.6.3 Simulation: the realistic answer
To find out what the design can actually detect, we simulate it. Rather than plug numbers into a formula, we generate thousands of artificial datasets that reproduce the real design—the continuous zero-sum share, six tasks, respondent-level clustering, and a moderator measured with realistic error—then run the exact model the analysis will use and count how often it detects each effect. The proportion detected is the power. Because no earlier study reports how large the moderation effect might be in this setting, we do not guess a single value: we sweep a range of true interaction sizes, from 0.5 to 3 pp per standard deviation of the moderator. The upper end of that range is anchored on Gilgen (2022), whose class-by-merit interaction in a comparable scholarship experiment is essentially null, which is our best available signal that real interaction effects here are small.
The simulation does two things. For the main effects, it simply confirms the benchmark: power is at or near 1.00 for a 6 pp effect at 1,500 respondents under both methods, so the continuous outcome changes nothing and the design is unambiguously powered.
For the interaction, it corrects the benchmark. Here the two methods diverge sharply, and the divergence is the whole point of running the simulation. For a 3 pp interaction at six tasks, the binary formula reports 38% power while the simulation reports essentially 100%. This is not a contradiction: the two numbers describe the same effect under different assumptions about outcome variability. The formula assumes the binary worst case; the simulation uses the real, much lower, variability of the share. The entire gap between 38% and 100% is the cost of that worst-case assumption—precisely the conservatism this subsection was meant to remove.
| N tasks | True interaction (pp/SD) | Interaction (sim) | Interaction (binary) | Gain |
|---|---|---|---|---|
| 3 | 0.5 | 0.475 | 0.039 | 0.436 |
| 4 | 0.5 | 0.555 | 0.041 | 0.514 |
| 5 | 0.5 | 0.755 | 0.044 | 0.711 |
| 6 | 0.5 | 0.745 | 0.046 | 0.699 |
| 3 | 1.0 | 0.980 | 0.058 | 0.922 |
| 4 | 1.0 | 0.995 | 0.065 | 0.930 |
| 5 | 1.0 | 1.000 | 0.072 | 0.928 |
| 6 | 1.0 | 1.000 | 0.079 | 0.921 |
| 3 | 1.5 | 1.000 | 0.084 | 0.916 |
| 4 | 1.5 | 1.000 | 0.099 | 0.901 |
| 5 | 1.5 | 1.000 | 0.113 | 0.887 |
| 6 | 1.5 | 1.000 | 0.127 | 0.873 |
| 3 | 2.0 | 1.000 | 0.118 | 0.882 |
| 4 | 2.0 | 1.000 | 0.143 | 0.857 |
| 5 | 2.0 | 1.000 | 0.169 | 0.831 |
| 6 | 2.0 | 1.000 | 0.194 | 0.806 |
| 3 | 2.5 | 1.000 | 0.161 | 0.839 |
| 4 | 2.5 | 1.000 | 0.200 | 0.800 |
| 5 | 2.5 | 1.000 | 0.239 | 0.761 |
| 6 | 2.5 | 1.000 | 0.277 | 0.723 |
| 3 | 3.0 | 1.000 | 0.212 | 0.788 |
| 4 | 3.0 | 1.000 | 0.268 | 0.732 |
| 5 | 3.0 | 1.000 | 0.323 | 0.677 |
| 6 | 3.0 | 1.000 | 0.376 | 0.624 |
Because the simulation gives realistic power, we can read off the smallest interaction the design can reliably detect. At 1,500 respondents and six tasks, that floor is about 0.7 pp per standard deviation of the moderator once the moderator’s measurement error is taken into account.
| N tasks | MDE, reliability = 0.8 (realistic) | MDE, reliability = 1.0 (benchmark) |
|---|---|---|
| 3 | 0.9 pp | 0.82 pp |
| 4 | 0.82 pp | 0.78 pp |
| 5 | 0.8 pp | 0.59 pp |
| 6 | 0.73 pp | 0.61 pp |
What this means for the moderation hypotheses is a matter of size. The design detects interactions of moderate size with ease: an effect of 1 pp per standard deviation is caught with power near or above 0.90 at every task count, and anything larger with near-certainty. It cannot reliably detect very small interactions: an effect of 0.5 pp per standard deviation falls below the 80% threshold at 1,500 respondents.
This sets a clear and honest expectation. The comparable evidence suggests that the true moderation effects in this domain may be small—possibly at or below the 0.7 pp floor. We therefore treat the moderation hypotheses with corresponding caution in this first cross-sectional wave: if we find no moderation, that null is consistent with the existing literature and not a symptom of an underpowered design; if we find moderation large enough to detect, it is a substantive result. Because the detectable floor shrinks as the sample grows, a larger later wave is the natural place to test these hypotheses as fully confirmatory. The main-effect hypotheses, by contrast, are firmly powered now.
3.7 Quality control
All checks below are pre-registered and used to flag rather than automatically exclude respondents: the confirmatory analysis is reported both with and without flagged cases (see Missing data).
3.7.1 Attention and comprehension checks
Two screening items are included. A comprehension check, placed after the practice task and before the six allocation tasks, asks respondents what their task will consist of, with the correct option (“distributing a scholarship amount between two applicants”) shown alongside distractors describing plausible misreadings, most importantly choosing a single applicant, which corresponds to the forced-choice format the design does not use. An attention check of the standard instructional-manipulation type, embedded among the post-task attitudinal items, instructs respondents to select a specific option (“disagree”) regardless of content. Respondents failing either are flagged.
3.7.2 Time filters
Total completion time and time on the conjoint module are recorded. Respondents whose completion time falls implausibly below the median are flagged as likely satisficing. Because the plausible lower bound depends on the observed distribution, the exact threshold is fixed after the pilot rather than imposed in advance, and pre-registered before the main study.
3.7.3 Response patterns
Two non-substantive patterns are monitored: invariant responding across the six tasks—most saliently leaving every slider at the default even split—and straight-lining across the attitudinal batteries used to construct the moderators. Because a fixed even split is also a substantively meaningful choice, these patterns are flagged for sensitivity analysis rather than used as automatic exclusions.