Conjoint Design Expert

SkillMedia

Designs conjoint and factorial-vignette experiments end to end — attribute architecture and randomization restrictions, effective-N power from the closed-form AMCE standard error, treatment realism, estimand choice among AMCE, marginal means, and AMIE, design variants such as forced choice versus rating, PAP tiers for conjoint flexibility, and the regression models and R packages that implement each. Use when the user is planning a conjoint, drafting or critiquing an attribute table, asking how many respondents or tasks are needed, asking whether to report AMCEs or marginal means, or asking how to test interactions. Reviewing an existing design goes to conjoint-diagnostics, cleaning the export to conjoint-cleaning.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Conjoint Design Expert skill

What this skill tells your AI

The instructions your AI receives, as published by scdenney/open-science-skills in codex/conjoint-design/SKILL.md and read by ahel’s review.

Instructions

Worked example (attribute table → power calculation → PAP tier assignment): see references/example.md.

1. Attribute Architecture

  • Orthogonality: Ensure every attribute is independent of every other attribute to allow for the estimation of causal effects for each component.
  • Randomization of Order: Order attributes randomly at the respondent level (not the task level) to prevent "primacy" or "recency" effects while avoiding the cognitive overload of finding information in different orders across tasks (Stantcheva 2023). A specific logical flow may override this if theoretically required.
  • D-Optimal Designs: Consider D-optimal or constrained randomization schemes rather than pure randomization. D-optimal designs choose the sets of administered conditions that maximize statistical power and may be preferable when the number of possible attribute combinations is large relative to the sample size (Auspurg & Hinz 2015; Stantcheva 2023).
  • Attribute Density: Monitor for respondent fatigue. Stefanelli and Lukac (2020) cite evidence that conjoint results remain stable with up to 10 attributes; Bansak et al. (2018) find that response quality does not degrade with up to 30 tasks on MTurk, and Bansak et al. (2021, "Beyond the Breaking Point") extend this to the attribute dimension, reporting stability at many attributes. These are the canonical sources for the task- and attribute-count claims respectively. Still evaluate whether the complexity of the levels increases cognitive load beyond the attribute count alone.
  • Nested/Constrained Randomization: Not all attributes need to be fully crossed. When ecological validity demands it, certain attribute levels can be linked or nested within other attributes (e.g., origin countries nested within policy domain). This is acceptable when: (a) the nesting is theoretically justified, (b) the primary attributes of interest remain fully independently randomized, and (c) the analyst acknowledges that nested attributes cannot be cleanly separated from their parent attribute. See Auspurg & Hinz (2015) on restricted randomization in factorial surveys.
  • Attribute-Level Restrictions: Implausible combinations can be excluded when they would confuse respondents or produce artifactual responses, but this is a judgment call, not a mandate. Eye-tracking evidence from Bansak & Jenke (2025) shows that odd (incongruent or nonsensical) attribute combinations have minimal, inconsistent effects on respondent attention, search, and choice, so decisions to include or exclude them should be driven primarily by statistical, substantive, and theoretical considerations (e.g., whether the randomization distribution should reflect a real-world target distribution per De la Cuesta, Egami & Imai 2022) rather than by assumed cognitive-burden concerns. Document all restrictions in the pre-analysis plan.
  • Medium-Level Specificity: Attribute levels should be concrete enough to be meaningful but not so specific that they introduce unintended confounds. Describe treatments at a "medium level of specificity" -- "fully described but not overly described" (Sniderman 2018). Avoid vague descriptions (e.g., "a policy that helps the economy") and overly narrow ones (e.g., "a $2.3B infrastructure bill for Route 95 in Pennsylvania").

2. Statistical Power and Error Logic

  • Effective N (N_eff): Calculate sample size based on (Respondents $\times$ Tasks $\times$ Profiles). Throughout this section, N_eff refers to this effective number of profile evaluations. However, respondents and tasks are not interchangeable -- adding respondents improves precision more than adding tasks per respondent due to within-respondent correlation. When in doubt, prioritize more respondents over more tasks (Stefanelli and Lukac 2020).
  • Closed-Form Formula: The standard error of an AMCE is approximately: SE = $\sqrt{\text{Var}(Y) \times L / N_{\text{eff}}}$, where $L$ is the number of levels for the attribute and N_eff is as defined above (Schuessler and Freitag 2020). This provides a quick diagnostic for whether precision is adequate.
  • Interaction Power: Estimating interaction effects requires approximately twice the sample size needed for main AMCEs, in the canonical balanced two-level-by-two-level case; the exact multiplier scales with the number of levels on each interacting attribute. The standard error of an interaction coefficient is approximately $\sqrt{2}$ times the SE of the corresponding main effect in that balanced case (Schuessler and Freitag 2020). Budget accordingly when interaction hypotheses are confirmatory.
  • Empirical AMCE Benchmarks: Typical AMCEs in published conjoint studies range from 0.02 to 0.10 (percentage-point changes in selection probability), with a median around 0.05 (Stefanelli and Lukac 2020). Very large AMCEs (> 0.15) are rare. Use these benchmarks when setting the smallest effect size of interest (SESOI) if no prior data are available.
  • Minimum Detectable Effect (MDE): Set the MDE based on the attribute with the highest number of levels, as this level will be the most difficult to estimate precisely. Report whether the MDE falls within the range of plausible AMCEs given prior literature.
  • Type S and Type M Errors: When power is low, beware of "Type S" (Sign) errors (getting the direction wrong) and "Type M" (Magnitude) errors (exaggerating the effect size). At 50% power for a true effect of d = 0.5, the probability of a Type S error is approximately 1/18, and the expected Type M error (exaggeration ratio) is approximately 1.5 (Lakens 2025, citing Gelman and Carlin 2014).
  • Low-N_eff Danger Zone (rule of thumb): Designs with fewer than ~3,000 effective profile evaluations are at high risk of being underpowered for detecting typical AMCE magnitudes (0.02–0.05). This threshold is a pragmatic rule of thumb derived from plugging median-AMCE benchmarks into the Schuessler and Freitag (2020) formula, not a research finding; adjust based on the expected AMCE magnitude, number of levels, and design. Below this threshold, conduct an explicit sensitivity analysis showing what effects can be detected.
  • Levels-Power Tradeoff: Each additional level for an attribute reduces the effective number of observations per level. As an illustration, going from 4 to 5 levels reduces per-level N by about 20%, with a corresponding precision loss (Schuessler and Freitag 2020). Only add levels when each is theoretically necessary.
  • Multiple Testing: Conjoint designs estimate many AMCEs simultaneously, which inflates the family-wise false-positive rate. Even under the sharp null of no effects, a typical conjoint pipeline produces at least one significant AMCE in more than 90% of experimental trials (Liu and Shiraito 2023); this mirrors the broader "garden of forking paths" problem (Gelman and Loken 2014) and the undisclosed-flexibility findings of Simmons, Nelson, and Simonsohn (2011), which motivate pre-specification and correction. Pre-specify a correction method: Bonferroni (most conservative, guards against false positives), Benjamini-Hochberg (controls FDR, most lenient), or adaptive shrinkage (balanced; preferred in Liu and Shiraito's simulations). Report both corrected and uncorrected results for confirmatory hypotheses.
  • Cohen's d Warning: Do not use Cohen's d benchmarks (small = 0.2, medium = 0.5, large = 0.8) to calibrate conjoint power analyses. AMCEs are measured in percentage-point changes in choice probability, not in standard deviation units. Translating between the two requires knowing Var(Y), which depends on the choice task structure.
  • Tools: Use the cjpowR R package (Freitag 2021) or the associated Shiny app for simulation-based power analysis. These allow specification of the number of attributes, levels, tasks, and profiles, and return power curves for main effects and interactions. For a general declare-diagnose-redesign workflow that couples the closed-form formula with design-based simulation across estimands, diagnosands, and assignment schemes, use the DeclareDesign framework (Blair, Cooper, Coppock, and Humphreys 2019). For interaction analysis, use FindIt (Egami and Imai 2019). For heterogeneity detection, use cjbart (Robinson and Duch 2024) or the Bayesian mixture-of-regularized-regressions approach in Goplerud, Imai, and Pashley (2025). For lexicographic preference ranking, use cjRank (Dill, Howlett, and Müller-Crepon 2024). For assumption-free tests of whether a factor matters at all, use CRTConjoint (Ham, Imai, and Janson 2024). For deploying an adaptive focal/context design, use the Docker container at github.com/dmolitor/adaptive-infra, with replication scaffolding at github.com/jennahgosciak/adaptive_conjoint (Gosciak, Molitor, and Lundberg 2026); standard survey platforms (Qualtrics) do not support continuous Thompson-sampling updates.
  • Compromise Power Analysis: When the respondent pool is fixed (e.g., hard-to-reach populations, elite samples), use a compromise power analysis that balances Type I and Type II error rates. An alpha > 0.05 may be defensible when it minimizes the combined error rate (Lakens 2025).

3. Treatment Validation and Realism

  • Identify the DGP Before Attributes: Before specifying attributes, articulate the data-generating process implied by the theory: which component of the compound profile bears the causal effect of interest, and under what assumptions about how respondents bundle or separate information. This aligns with the broader project principle of DGP identification and with the "Why → If-Then" funnel developed in the hypothesis-building skill.
  • Experimental vs. Mundane Realism: Distinguish between experimental realism (does the task engage respondents psychologically?) and mundane realism (does the task resemble real-world decisions?). Mundane realism is neither necessary nor sufficient for validity -- what matters is that the treatment creates the intended psychological state (Druckman 2022, citing Aronson and Carlsmith 1968). For conjoints, tabular displays may lack mundane realism but achieve experimental realism if respondents attend to and process the information.
  • Attention and Salience as Generalizability Levers: Even when internal validity is secured, survey-experimental effects may amplify or reverse real-world effects because the survey environment compresses the consideration set (attention) and distorts the relative weights on attributes (salience). Fu and Li (2024) formalize these two mechanisms: consideration-set compression tends to inflate AMCE magnitudes, and context-dependent salience can flip effect signs. Audit a conjoint against both: does the table force attention to attributes respondents would ignore in the real-world analog, and are the attributes displayed with relative salience that mirrors the target environment?
  • Information Availability, Access, and Processing: Validate that respondents (1) have access to the attribute information (can they see it?), (2) attend to it (do they read it?), and (3) process it as intended (do they interpret it the way the researcher assumes?). Attention checks and comprehension probes address conditions 1-2; pilot studies and cognitive interviews address condition 3 (Druckman 2022).
  • Names-as-Cues Warning: When attributes include proper names, cultural referents, or country names, these may carry unintended associations beyond the dimension of interest. Pilot-test whether respondents associate additional meanings with the selected names or labels.
  • Pretreatment Mock Vignette: Consider presenting respondents with a non-experimental practice vignette before the conjoint block to familiarize them with the task format. This reduces learning effects across early tasks.
  • Repeated Task for IRR Estimation: At design stage, plan to repeat the first conjoint task at the end of the block with left/right profile order reversed. This provides a direct estimate of intra-respondent reliability (IRR) for measurement error correction (Clayton et al. 2023), who report ~75% intra-respondent agreement on identical tasks across eight replicated conjoint studies. Clayton et al. find respondents generally do not detect the repetition at rates that bias the IRR estimate. The repeated task costs one additional task per block but enables bias-corrected AMCEs and marginal means via the projoint R package. This is a design decision — it cannot be retrofitted after data collection.
  • Assumption-Free Factor Tests: Complement the standard Hainmueller, Hopkins, and Yamamoto (2014) diagnostics with the conditional randomization test in CRTConjoint (Ham, Imai, and Janson 2024), which provides an assumption-free test of whether a factor of interest matters in any way given the other factors, and tests for profile-order, carryover, and fatigue effects. This is especially valuable when AMCE-based confidence intervals are narrow and contain zero — a narrow AMCE CI implies a weak marginal effect, not necessarily a weak total causal effect.

4. Estimating Effects

  • Reference Categories: Clearly identify the "baseline" or "reference" level for every attribute.
  • SESOI in AMCE Terms: For every confirmatory hypothesis, state the smallest meaningful AMCE -- the smallest percentage-point change in choice probability that would be theoretically or practically significant. If 2 percentage points is considered trivially small, state this as the lower bound. Justify the SESOI based on prior conjoint studies, theoretical significance, or policy relevance (Lakens 2025).
  • Average Marginal Component Effects (AMCE): Frame results as the average change in the probability of being selected when an attribute changes from the reference level to the level of interest. Note that the AMCE "critically relies upon the distribution of the other attributes" used for averaging (De la Cuesta, Egami, and Imai 2022) -- this is typically the uniform distribution imposed by the randomization, which may not match real-world attribute distributions. When the target of inference is a specific real-world or counterfactual distribution, use design-based weighting (e.g., via factorEx) rather than the default uniform AMCE. Define the estimand explicitly — unit-specific quantity, target population, and aggregation — before selecting an estimator, per the estimand-first framework of Lundberg, Johnson, and Stewart (2021).
  • AMCE Interpretation Guard-Rail (no majoritarian claims): The AMCE does not identify the share of respondents who prefer a given feature. It aggregates both the direction and the intensity of preferences, so statements of the form "voters prefer X" or "Americans prefer Y" are not supported by a statistically significant AMCE unless additional (strong) assumptions hold, e.g., uncorrelated direction and intensity (Abramson, Koçak, and Magazinnik 2022). Ganter (2023) sharpens the same point: AMCE identifies a population-level effect on choice probability, not a parameter of the underlying preference distribution, and conflating the two has produced widespread interpretive overreach. Valid interpretations include: change in expected vote share over the experimenter-defined contest distribution, mapping to a Borda-rule winner, or an average of individual ideal points. If the substantive target is a majoritarian or electoral claim, consider the bounding method and structural interpretations in Abramson et al. (2022), Ganter (2023), and the outcome-variable/estimand contingency analyses in Treger (2025).
  • Marginal Means (MMs): In addition to AMCEs, report marginal means -- the model-predicted probability that a profile is selected when a given attribute level is shown, averaged over all other attributes. MMs provide absolute levels of support rather than relative differences, and are particularly useful for subgroup comparisons (Leeper, Hobolt, and Tilley 2020).
  • Interaction Effects via AMIE: Do not test for attribute-by-attribute interactions by adding product terms to a dummy-coded AMCE regression -- such interaction coefficients are baseline-dependent artifacts that change when the reference category changes (Egami and Imai 2019). Instead, use the Average Marginal Interaction Effect (AMIE), defined as the additional effect of an attribute combination beyond the sum of its separate AMEs: AMIE(a, b) = ACE(a, b) − AME(a) − AME(b). The AMIE is invariant to the choice of baseline condition and is the only interaction estimand interpretable across coding schemes. Estimate AMIEs using penalized ANOVA via the FindIt R package (CausalANOVA()), which simultaneously handles the high-dimensionality problem (even a modest conjoint with 5 attributes and 4 levels each generates 100+ interaction parameters) through regularization that shrinks weak interactions toward zero and collapses adjacent levels with similar effects (Egami and Imai 2019). Note: the AMIE framework applies to interactions between randomized conjoint attributes, not to interactions between attributes and non-randomized respondent characteristics (subgroup moderators).
  • Heterogeneity Detection via BART or Mixture Models: To detect treatment effect heterogeneity across respondents without pre-specifying moderators, two complementary approaches are available.
    • BART (cjbart): Robinson and Duch (2024) fit a probit BART model to estimate Individual-level Marginal Component Effects (IMCEs) -- each respondent's predicted effect for each attribute level. The method introduces a three-level estimand hierarchy: OMCE (observation-level) → IMCE (individual-level, averaged across tasks) → AMCE (population-level, averaged across respondents). The IMCE distribution is the primary heterogeneity diagnostic: a tight, normal distribution centered on the AMCE suggests homogeneous effects; multimodal, skewed, or widely dispersed distributions (especially spanning both sides of zero) indicate substantive heterogeneity. Use het_vimp() to identify which respondent covariates most strongly partition the IMCE distribution via random forest variable importance scores.
    • Bayesian mixture of regularized regressions: Goplerud, Imai, and Pashley (2025) propose a finite mixture of regularized regressions that groups respondents exhibiting similar treatment-effect patterns and directly models cluster membership with covariates. This is preferable when interpretable clusters of respondents and their moderator-driven membership are of scientific interest, and when hierarchical interaction structure should be enforced.
    • Heterogeneity detection with either tool is exploratory by default and should be labeled as such unless moderators and directions were pre-registered (Robinson and Duch 2024).
  • Profile-Context Heterogeneity (the GML estimand): BART, FactorHet, and CRT all reason about heterogeneity across respondents. Gosciak, Molitor, and Lundberg (2026) introduce a complementary estimand for heterogeneity across the profile context space: $\theta(\vec{x}) = \Pr(\text{choose focal} = A \mid \text{context} = \vec{x})$ and the two summary parameters $\theta(\vec{x}{Max}), \theta(\vec{x}{Min})$ that bound the range a focal-attribute effect can take across all combinations of the other attributes. The estimand reframes the conjoint by designating one attribute as focal (the causal quantity of interest) and the rest as context (a vector $\vec{x}$ that conditions the focal effect). Use this when the theoretical claim is configural — "the effect of X depends on the full bundle of who the immigrant/candidate/job applicant is" — rather than additive. Two-way AMIEs and the CRT can detect that heterogeneity exists but cannot localize where it lives; $\theta(\vec{x}{Max})$ and $\theta(\vec{x}{Min})$ pin down the cells that maximize and minimize the focal effect, with all other attributes held jointly fixed at $\vec{x}$. Even without an adaptive design, the estimand is computable on any standard conjoint by reading off conditional cell means, at the cost of thin cells in rare combinations of context attributes (see §5 Adaptive Randomization variant for the data-collection design that targets this estimand directly).
  • Nested Marginal Means for Lexicographic Preferences: When the theory predicts that respondents may have categorical (lexicographic) preferences -- always vetoing profiles with a given attribute level regardless of other attributes -- standard AMCEs and marginal means can be misleading because co-occurrence rates across task pairs artificially inflate or deflate estimates. Use nested marginal means (Dill, Howlett, and Müller-Crepon 2024): (1) identify the attribute with the highest predictive power (Rank 1), (2) compute marginal means for remaining attributes restricted to tasks where the Rank 1 attribute does not vary between profiles, (3) iterate to rank remaining attributes. This reveals whether attributes that appear unimportant are genuinely so or are simply dominated by a higher-priority veto attribute. Implemented in the cjRank R package.
  • Equivalence Testing for Null Predictions: When the hypothesis predicts that an attribute has no meaningful effect (e.g., a manipulation check attribute), use the TOST equivalence testing procedure with a pre-specified equivalence range in raw AMCE units rather than simply reporting a non-significant p-value (Lakens 2025).

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
55
Forks
3
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
conjoint-design
Source
github.com/scdenney/open-science-skills