BrainChallengeEducation

Stroop Test

Run a proper two-block Stroop test and get your interference effect in milliseconds. Congruent and incongruent blocks, keyboard timing, mean reaction time and accuracy for each block.

Respond to the INK COLOUR, never the word.

If the word GREEN is printed in blue, the answer is BLUE.

Trials per block
(40 trials total)

You’ll do two blocks back to back. In the first, the word and the ink match. In the second, they never match. Name the ink colour as fast as you can.

or press Space

What this test actually measures

A Stroop test does not produce one score. It produces a difference between two scores. You run a congruent block (the word RED printed in red ink) and an incongruent block (the word RED printed in blue ink), naming the ink colour in both. Subtract the congruent mean reaction time from the incongruent mean, and the leftover milliseconds are the interference — the cost of suppressing a response you did not want to make.

That subtraction is the whole point. Your raw reaction time is dominated by things that have nothing to do with attention: your display’s refresh rate, your keyboard’s scan latency, your age, how much coffee you have had, whether you are using a trackpad. Those costs land on both blocks roughly equally, so they cancel out of the difference. What survives the subtraction is closer to a real cognitive quantity.

Two other conventions matter. First, mean reaction time is computed over correct trials only — an error trial measures a wrong decision process, not the one under study, and errors are often suspiciously fast because they are the un-inhibited reading response slipping through. Second, accuracy is reported alongside RT, because someone can trade one for the other: answer slowly and accuracy stays high with a small interference effect; answer fast and errors pile up in the incongruent block instead. Neither number alone tells you anything.

Stroop’s 1935 experiment — what he actually found

John Ridley Stroop published Studies of interference in serial verbal reactions in the Journal of Experimental Psychology in 1935, as part of his doctoral work at George Peabody College. The phenomenon was not new — Wilhelm Wundt’s lab, and later James McKeen Cattell in 1886, had already noted that naming a colour takes longer than reading a colour word — but Stroop was the first to isolate the interference cleanly and quantify it.

His paper ran three experiments, and the third is the one most people forget:

ExperimentTaskResult
1Read the word, ignoring the conflicting ink colourEssentially no cost — about 2.3 seconds slower across a whole 100-item sheet, which Stroop treated as negligible
2Name the ink colour, ignoring the conflicting wordA 74% slowdown — roughly 47 seconds added to a 100-item sheet compared with naming solid colour patches
3The same colour-naming task, after eight days of practiceInterference shrank substantially with training — but rebounded once practice stopped, and practice at colour naming then increased interference on the word-reading task

The headline finding is the asymmetry, not the slowdown. Colour words wreck colour naming; colours barely touch word reading. Any explanation of the Stroop effect has to account for that one-directionality, and that constraint is what has kept the task alive for ninety years. Note also that Stroop measured total time to complete a printed card, not per-trial reaction times — the trial-by-trial millisecond version used here (and in this test) came later, with computerised presentation.

Why the interference happens — the competing accounts

There is no settled single answer. These are the main explanations, each of which predicts the asymmetry for a different reason:

  • Automaticity of reading: the oldest account. For a literate adult, reading is obligatory and requires no intention — the word’s meaning is retrieved whether you want it or not, and then has to be suppressed. Colour naming, by contrast, is deliberate and effortful. Supporting evidence: the effect is weak or absent in pre-readers and grows as reading fluency develops through primary school, peaking around ages 7–12 in relative terms. The weakness of the account is that “automatic” turns out not to be all-or-nothing — the effect shrinks when the word is made harder to see, which a truly obligatory process should not permit.
  • Speed of processing (relative speed): the word is simply processed faster than the ink colour, so its response arrives first and occupies the output channel. Whichever pathway finishes first interferes with the slower one. This neatly explains the asymmetry without invoking automaticity, but it struggles with experiments that speed up colour processing or slow down word processing and still find interference in the same direction.
  • Parallel distributed processing (Cohen, Dunbar & McClelland, 1990): the account that currently does the most work. Word reading and colour naming are two pathways in a connectionist network, and the word-reading pathway has stronger connection weights because of far more lifetime practice. A task-demand unit biases processing toward the relevant pathway, but that top-down bias has to overcome the stronger pathway’s head start. Interference emerges from the relative strength of the two pathways, not from a categorical automatic/controlled distinction — which explains why the effect is graded, why it shrinks with practice at colour naming (Stroop’s own Experiment 3), and why it is sensitive to stimulus quality.
  • Where the conflict lives: a separate debate asks whether interference arises at the semantic level (two meanings competing) or the response level (two spoken or keyed responses competing). Evidence for both exists: interference shrinks but does not vanish when the distractor word is a colour that is not one of the response options (semantic but non-response competition), which suggests the two contribute separately.

The cognitive constructs involved

The Stroop task is usually described as an index of executive function. That is a family of abilities rather than one thing, and the task touches several of them:

  • Selective attention: holding the relevant feature (ink) in focus while a competing feature (word) occupies the same physical location. Unlike most attention tasks, you cannot solve this one by looking somewhere else — the two dimensions are fused in a single object.
  • Inhibitory control (response inhibition): suppressing the prepotent, over-learned reading response in favour of the weaker, instructed one. In Miyake’s influential three-factor model of executive function, Stroop loads primarily on the inhibition factor.
  • Task-set maintenance: keeping “name the ink” active in working memory across dozens of trials while the stimulus itself keeps arguing for the other rule. Lapses here show up as occasional very slow trials or clusters of errors rather than a uniform slowdown.
  • Cognitive flexibility: engaged mainly in switching versions of the task, where the rule alternates between reading and colour naming. The fixed-rule version run here taxes it far less — a point worth being precise about, since flexibility is often claimed for the standard task without justification.
  • Processing speed: a confound rather than a construct. Slower responders often show larger raw interference simply because all their reaction times are scaled up. This is why researchers sometimes report a proportional score (interference divided by congruent RT) alongside the raw millisecond difference.

On reliability: the Stroop effect is one of the most robust findings in psychology at the group level — it replicates almost every time it is run. Its individual-difference reliability is much weaker. Because it is a difference score between two noisy measurements, test–retest correlations for a person’s interference magnitude are often modest. This is the “reliability paradox” described by Hedge, Powell and Sumner (2018): tasks designed to minimise between-person variance make poor individual-difference instruments. Practically: expect your own number to bounce around between runs, and do not read much into a single result.

Research and clinical use — and what this page is not

Standardised Stroop instruments exist and are used in neuropsychological assessment — the Golden version (1978), the Comalli–Kaplan version, the Delis–Kaplan Executive Function System (D-KEFS) Color-Word Interference subtest, and the Victoria Stroop Test among them. They are administered by trained clinicians, usually with spoken responses on printed cards, scored against age- and education-stratified normative samples, and interpreted only as one input among many.

In research, Stroop interference has been studied in relation to ADHD, frontal-lobe injury, schizophrenia, depression, dementia, and the effects of sleep loss, alcohol, and stress. The task also has a substantial neuroimaging literature: incongruent trials reliably increase activity in the dorsal anterior cingulate cortex — associated with conflict monitoring — and in dorsolateral prefrontal cortex, associated with implementing top-down control. It is a workhorse for probing that control network.

This web version is not a diagnostic instrument. It has no normative sample, no clinician, no control over your screen, input device, lighting, or whether anyone was talking to you halfway through. Browser timing adds variable latency, and 20 trials per block is a fraction of what a research protocol uses. It cannot detect, screen for, rule out, or measure the severity of ADHD, dementia, brain injury, or any other condition. If you are concerned about your attention or memory, that is a conversation for a clinician, not a web page.

Known variants of the task

  • Reverse Stroop: respond to the word while ignoring the ink. Interference is much smaller or absent when the response is spoken — this is Stroop’s Experiment 1. It reappears, however, when the response is made by pointing to a colour patch rather than speaking, which is one of the better arguments that response modality, not just stimulus processing, shapes where conflict arises.
  • Emotional Stroop: name the ink colour of emotionally loaded words (for example threat words for anxious participants, body-shape words in eating-disorder research). Slowing on those words is taken as an attentional-bias index. Importantly, this is mechanistically different from the classic effect: there is no competing colour response, so the slowdown reflects attentional capture by salient content rather than response conflict. The two should not be treated as the same measure.
  • Numerical Stroop: two digits differ in both numeric value and physical size (a small 8 beside a large 3). Judging physical size is slowed by an incongruent numeric value and vice versa, showing that magnitude is extracted automatically. Used heavily in numerical-cognition and dyscalculia research.
  • Spatial Stroop: the word ABOVE presented below fixation, or an arrow pointing left appearing on the right. Closely related to the Simon effect, and useful because it removes reading fluency as a factor.
  • Counting / day–night Stroop: report how many items are shown when the items are themselves digits, or say “day” to a moon picture. The day–night version is widely used with young children who cannot yet read well enough for the colour-word task.
  • Proportion-congruent manipulations: not a stimulus variant but a design one. If most trials in a block are incongruent, interference shrinks — participants strategically down-weight the word pathway. This is why the two blocks in this test are kept pure (100% congruent, then 100% incongruent), and also why blocked designs like this one typically produce somewhat larger effects than randomly mixed designs.