RSS Amplifier

The Pennsylvania Heretic · Mar 1, 2025

The CRIE: Correlations Between Intelligence Ratings and Intelligence Test Questions

0
Sign in to vote or save

Chapin Lenthall-Cleary, Cole Gaboriault · The Pennsylvania Heretic

Abstract

This paper examines the correlations between scores on a prototype intelligence test and ratings of intelligence. We found that most questions had negligible correlations with ratings of intelligence, but several had decent correlations: two questions had hybrid taus of -.435 and -.405 with the mean of Chapin's and Cole's ratings, and the exam as a whole had a hybrid tau of -.429 with our ratings. (Negative correlations indicate higher seeming intelligence correlating with a better score, since lower scores are better.) We calculated a 3.0σ significance of question 9 being correlated with this mean rating, along with 2.3σ for question 12 and 2.4σ for the whole exam score. These results suggest that ratings of intelligence, at least from some people, partly correlate with performance on certain cognitive tasks.

IMPORTANT: If you wish to take this exam, jokingly named the CRIE, do so before reading the rest of this article.

Methods

The test linked at the bottom of this paper was administered to 20 people, 15 of whom also received seeming intelligence ratings from at least 2 raters (that procedure is described here). Most of Chapin's and Cole's seeming intelligence ratings were given prior to seeing CRIE scores. Test-takers were encouraged to ask any clarifying questions, including any strictly-knowledge questions (since the test was intended to not be a knowledge test). Some participants were also given a social section involving guessing their and others' scores. This was administered after the exam, and time spent on it was not counted toward the exam time.

Chapin designed the exam, so we were able to give Cole a regular score (though the version of the exam he took was slightly different). Cole's score was included in our analysis. We also estimated what Chapin's score would've been. Since this is only a rough estimate, Chapin's proxy-score was excluded from our analysis.

To calculate significance, we performed a bootstrap test. We generated 1 million random pairs of fake seeming intelligence and score data sets by resampling separately from the Combined Unweighted Chapin+Cole ratings and the relevant scores (with replacement) and calculated how often they produced a hybrid tau magnitude greater than that between the actual ratings and scores (Q9:2346/M;Q12:17367/M,S:14037/M).

Notable Questions

The two questions best-correlated with seeming intelligence are below:

9.

Does the above graph show temperature consistently increasing as the number of pirates decreases? Briefly explain why or why not. max 2pts

12. List all positive integers between 1 and 50 (inclusive) that satisfy all of the following:

Is a multiple of the length of their digits (78 becomes "seveneight", so is 10 long and 3 becomes "three", so is 5 long).

Differs from the nearest square number by an even number.

.5pts for each missing answer; .2 pts for each incorrect answer

I do not fully understand why these questions performed significantly better than the other questions. Question 9 involves noticing the flipped x-axis values (35000 and 45000), plus the ability to read a graph, which the vast majority of, likely all, CRIE-takers have been taught to do. Question 12 involves checking which numbers satisfy both constraints.

As far as I'm aware, there's no analytical shortcut to question 12: it's necessary to check every (or every-other-ish) number. This is not true for question 11:

11. List all possible sequences of 4 english letters that satisfy all of the following:

Each letter is strictly later in the alphabet than the one before it.

The first letter is as many positions away from "a" in the alphabet as the last one is from "z".

There are two letters in the alphabet between the second and third letters.

The second letter is twice as far in the alphabet as the first.

The third letter isn't "k" or "q".

.5pts for each missing answer; .2 pts for each incorrect answer

This question can be expressed as a system of equations and inequalities, which, when solved, yield an answer without any sort of guess-and-check. I would've expected that questions that are more amenable to mathematical analysis would better-correlate with seeming intelligence; this was not the case here. This is puzzling; our only theory on the poor performance of question 11 is that it's very vulnerable to careless errors causing completely wrong answers, whereas those errors tend to cause only partly wrong answers for question 12.

Results

As mentioned above, questions 9 and 12 are the best-correlated with nearly every rater's seeming intelligence ratings. 13 rescored is nearly as well-correlated, and 14 has some correlation. The total score is also correlated with the combined ratings comparably to the correlations with questions 9 and 12. Interestingly, while most of the correlations with the combined ratings also exist with Chapin's ratings, many of the other correlations disappear for other raters, but correlations with questions 9 and 12 persist for every rater. Cole's ratings are generally the weakest-correlated with the CRIE, but show a correlation with time in the expected direction, which other raters do not.

Hybrid Tau values of CRIE questions against seeming intelligence ratings. Combined UW Honest is an unweighted mean of all raters, excluding one whom we suspect gave dishonest ratings (most were identical).

Pearson's r values.

Weighted Hybrid Tau values.

Conclusions

If we assume that our combined seeming intelligence ratings accurately measure intelligence (which is a stronger claim than our core contention that they partly measure it), then these results suggest that the questions given range from poor to decent at testing intelligence. While this suggests that none are particularly suitable for an intelligence test, it points towards considerations for developing better questions.

If we're being skeptical toward our contention, this shows that the ratings of intelligence collected are somewhat correlated with performance on certain cognitive tasks. As with inter-rater reliability, this is a necessary and promising but insufficient result for that contention. This doesn't provide a definitive answer on the mechanism of the correlation: it's possible that there's some reason for the correlation other than intelligence being both partly perceived and partly measured. To give one possibility (without suggesting this is correct), it's possible that mathematical knowledge would influence both scores on some questions and perceptions of intelligence (even after accounting for the effect of intelligence upon mathematical knowledge). This does, however (with the caveats of small sample size and only moderately strong correlations), rule out ratings of intelligence being purely noise. While the question of whether ratings of intelligence measure intelligence is only weakly hinted at by these results (and depends upon a definition of intelligence), those ratings do, albeit very imperfectly, measure some capability.

A note for those curious: I tested o3-mini-high on the most difficult problems on the CRIE (11-14). It solved all except Q11. Even with generous prompting explaining this, it failed to treat A as the 1st letter, not the 0th (so it said that C, not D is twice as far in the alphabet as B, which would be correct if A were the 0th letter, which it isn't). On Q13, the LLM correctly suggested an additional issue with the argument that I hadn't considered (namely knowledge, by that definition, not being closed under implication).

Links

If you want to supply data for our research, please fill out this form.

If you wish to perform your own analysis, you can view our data here, partially obfuscated to protect anonymity. If you wish to perform any analysis that would require raw data, contact me at chapinalc@gmail.com.

You can view the full text of the CRIE here.

You can view the full text of the CRIE, annotated with answers, here.

Read the original on pennheretic.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.