In discussing ideas for large-scale conjoint I realized I probably jumped in at the deep-end for anyone who only has a passing familiarity with conjoint analysis and how it is designed and how it works.
This is then a ‘filler’ to give a background to what conjoint analysis is and some of the challenges this gives anything around large-scale conjoint, including current rules of thumb on conjoint design.
Conjoint analysis starts from the principle of breaking a product or service into attributes and levels. Th combination of a set of attributes with different levels creates a product profile. Then, in a choice task, respondents choose between different product profiles to indicate which one they would buy or prefer (possibly with a ‘none of these’ option). We analyze these choices to then work back to estimate the value or utility score for each of the different levels, and we can then model preference for different potential products defined using the attributes and levels.
What conjoint analysis does is solves how to choose the product profiles and to collect the choices, so as to be able to estimate the utility value (part-worth) that best explains the ‘value’ of each of the levels.
If you take a simple design with 4 attributes with 3 levels. Then there are 3 x 3 x 3 x 3 = 81 potential products that can be created by combining the levels of each attribute, and 3+3+3+3 = 12 utility scores or parameters or part-worths that need to be estimated - one for each level.
As the number of attributes, or the number of levels increases, the potential number of product profiles made from combining the attributes and levels increases multiplicatively, and the number of parameters to be estimated increases linearly.
When the potential product combinations increases into the hundreds and then thousands, the first problem that conjoint solves is how do you choose which products to test in the choice tasks? And how many choice tasks should you ask?
Basic choice-based conjoint analysis uses two approaches to dealing with the potential number of product profiles.
Firstly, it uses a fractional factorial design to reduce the number of profiles to be shown. A full factorial design is showing every possible combination, a fractional factorial approach using a carefully chosen selection of potential combinations to maximize the amount of information that can be extracted statistically, while minimizing the number of combinations that need to be tested.
The basic idea is to select combinations that are different from each other while also maintaining level balance (all levels seen the same number of times) and orthogonality (pairs of levels appear the same number of times).
For instance, if you have a phone with 64GB memory, 256GB storage at $200, then comparing this to a phone also with 64GB, 256GB storage at $250 only tests one variable - the price. But if the comparison was with a phone with 64GB memory, 256GB storage at $200 and a phone with 32GB, 512GB storage and $250 then there are three potential differences that can be compared.
A carefully designed fractional factorial design finds the smallest set of options that can be analysed to estimate all the key parameters. This is known as experimental design and is common in several fields.
For instance, instead of testing all 81 variations for a 3x3x3x3 design, an efficient design can use 9 combinations (Taguchi L9 - 3^4) to extract the main-effects for each level.
However, as the number of attributes increases, even an efficient design becomes too large for one person. So conjoint analysis uses a second method. The fractional factorial design is spread across the set of respondents, so each respondent only sees to a subset of the full design - for instance only getting 8 to 12 choice tasks to avoid overburdening.
The full design is still present, but it is spread across the sample as a whole.
For analysis, all the choice tasks are analyzed together as if they are all independent. If we have 100 respondents each doing 8 choice tasks, that is taken as 100 x 8 = 800 choices to be analyzed.
Originally this was done with logistic regression, which would only produce a total sample level model. However, most analysts would now use Hierarchical Bayes analysis which allows for individual level preferences to be imputed.
This technique of pooling choices across respondents is what allows for parameters to be estimated based on the total number of choices observed, not just the total number of respondents. In principle, you could use 1 choice across 800 people or 8 choices across 100 people. Researchers choose a balance because more choices per individual are needed in order for HB to derive individual-level utility scores.
The third element of boosting the statistical observations for estimating parameters is to look at the implicit comparison pairs that are made in a choice task.
If the choice is between A or B and A is chosen, then we have one ‘choice-pair’ that has been observed.
But if the choice is between A, B, C or D and A is chosen, then there are actually three choice-pairs being observed. It gives us the implicit preferences A>B, A>C and A>D.
In other words, the form of the choice-task also affects the amount of information we can extract.
For example, in MaxDiff, where best and worst are asked, adding the ‘worst’ question, creates more choice information. If A is best and D is worst, then we have A>B, A>C, A>D and B>D, C>D - so now five choice pairs. If we ask for a full ranking of the four elements we would complete the set with B>C.
The type and size of the choice task - the number of profiles and what choices we ask can create and add to the information available for analysis. If filters or sorts are added these add data about the choice preferences. Similarly, if an attribute has a pre-defined order (good, better, best), this also adds to the choice information available by adding constraints and preferences to the choice set.
With lots of attributes and lots of parameters, one raw option is just lots more sample. But the toolkit also includes fractional designs, sample size, the type of choice-task itself and applicable a-priori data. And where attributes are continuous or have lots of potential levels (e.g. price), options like interpolation or top-boxing may help reduce the number of levels and parameters that need to be estimated.
The biggest issue with lots of attributes and levels and lots of parameters is that the product profiles become increasingly complex. Visually, when making a choice, it becomes impossible to actually consider everything at the same time.
The design decision to focus on 5-7 attributes only is often for the problem of overwhelming participants, and not necessarily for statistical reasons. For a classic conjoint design with too many attributes being shown at one time, the comparisons and judgements become overwhelming.
Instead human decision-making looks for simplifications, based on shortcuts or short-circuiting, such as only fixing on price or a fixed set of attributes instead of the whole set. The problem this creates is that when decisions are short-circuited, for instance never choosing an item over a budget limit, this can potentially lead to misleading estimations for other attributes because selections conditioned by price in the first instance can then lead to other attributes being ignored or mis-selected due to the dominance of the short-cut element.
However, at this point we run into a contradiction.
Whenreal decision-making in e-commerce systems, or in-store, continually shows many more attributes in play and much more content to be considered. Online shopping systems and catalogues show lots more information with more variety and visual and textual noise than is ever shown on-screen in a conjoint exercise.
In areas, like software packages, where a large run of features is often given, care is taken on the presentation of the features to help purchasers focus on the features they determine as important.
Richer content also comes with a focus on layout, design and display, and adding variety makes the search and the choice more interesting.
By contrast, conjoint looks boring, and with a limited set of levels it looks and feels very repetitive and like a game of spot-the-difference.
Secondly, shortcuts and short-circuiting are real choice behavior with a real impact on purchasing and decision making. In numerical terms, a shortcut means other parameters can be treated as if they have no influence - they should have a zero-beta.
With a larger design is it enough to only estimate the important items driving the choice, and to allow for a lot of unattended or non-contributory items?
The counter-argument is that if you have items that have no contribution to the choice, why would you include them in the conjoint? It would cleaner to remove non-contributing attributes and levels, so as to make the actual conjoint design tighter and more focused. But if this comes at a cost of a smaller and tighter conjoint slowly looks less and less real is this the right way to go? Is realism important or valuable?
In the early days of conjoint and conjoint-type designs there used to be a lot more variation in terms of types or flavors of conjoint analysis. Sawtooth’s original Adaptive Conjoint Design was targeted at estimating 20-30 attributes using a mixture of pre-ranking and partial choices. Adaptive designs still exist in Adaptive Choice-based Conjoint, and conjoint variations like menu-based conjoint still appear.
But by-and-large, most conjoint now is choice-based conjoint which emerged from the econometrics field as stated-preference analysis or discrete choice estimation. The vociferously held view from the econometric community was that only choices that reflected purchasing mattered.
If we watch how people make decisions on e-commerce systems, the final purchase is not the only choice we can observe. Individuals add likes and dislikes, click on items they consider, add items to a basket or wish list and use features like sorting and filtering to help come to a decision. The final ‘this is what I would buy’ is generally part of a much longer process of choosing and evaluating.
These choices - such as likes and dislikes - can be used for choice-analysis. Secondly, the larger choice sets that are typically seen in e-commerce systems - which individuals potentially working through 50 or more items per page, on screen, increases the choice information available.
In a small conjoint with four items to choose from we get 3 choice-pairs. If a product is chosen from one of 50, then we have 49 choice pairs. But if we ask for like or dislike, or some form of selection or rejection we start to get a much richer choice data set than just the simple choose one from four.
However, for these choices to work, we also need enough variation in the products shown. We can still use fractional designs for efficiency, but we can also add variation in the form of textual noise and variety, and accept and look for short-circuiting behavior and zero-value parameters.
Lastly, we have to ask, do we need to collect parameters for all of the items? Often decision making will be based on the top two or three items. Everything else is an also ran.
If we know the preference order for the levels in an attribute, can we simplify the attribute to top-three plus other, or use top-middle-lowest, or good-better-best when considering the attribute design.
Secondly, one of my hypotheses is that selection and rejection are two complimentary models, not necessarily a single combined model. Items that contribute to a rejection model, do not necessarily contribute to a model that purely estimates drivers of selection. Conjoint combines them as a single preference model, which then requires all the paramaters to be estimated.
The outstanding question though, is whether we have the tools and methods to actually analyze large-scale choice data, and what would it look like? I look at the neural networks used for LLMs, which are not dissimilar to our little choice models at a very simple level - just way bigger and way deeper. If you tip all the real purchases on Amazon, or all the real hotel bookings into an LLM, would it have all the utilities for all the products hidden within its parameter set?
Having said all this, conjoint is extremely well-defined and well-grounded, with a robust history, and proven processes and methods. So the final question is whether there is actually any good reason to try to do it differently?
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.