Today I did all 31 released Evaluate questions. Each one asks some version of: what extra information would most help you judge whether an argument works? After working through all of them, I noticed five repeatable patterns that make these questions much more predictable.
1. Use the Two-Outcome Test
Evaluate questions are about finding information that can pull an argument in opposite directions depending on the answer. In practice, the right choice usually gives you two clear outcomes that meaningfully change how strong the conclusion looks.
When the issue is binary, test both sides and compare the impact:
- Are clouds a good indicator of rain coming? If yes, bring the umbrella. If no, maybe leave it.
- Do you mind getting soaked without an umbrella? If yes, bring it. If no, you can go without.
- Do you consider carrying an umbrella an annoyance? If yes, maybe skip it. If no, that favors bringing it.
The correct answer is the one where different plausible answers create meaningful shifts in confidence. If nothing really changes, that option is usually not doing true Evaluate work.
2. More than Two Answers? Check the Extremes
Not every Evaluate question is a clean yes/no setup. Some use percentages, probabilities, or money values. In those cases, push the answer choice to extreme ends and see how the conclusion reacts.
- For a percentage, imagine 0% versus 100%.
- For profit values, imagine a penny versus a trillion dollars.
If one extreme would strongly strengthen or weaken the conclusion, you are likely looking at the right Evaluate lever. A quick check I use is: if I were the author, would I care a lot whether this number goes up or down? If yes, keep it in play. If no, discard it even if it sounds topical.
3. When the Wrong Answer Sounds Right
The best wrong answers usually stay close to the topic while failing to affect the core reasoning. They feel relevant, but they do not pass the two-outcome test.
For example, if the argument is about whether a product is worth buying now, a choice about what the product cost in the past can sound useful but may not actually change the present-value judgment. Likewise, if the argument is about technical feasibility, an answer about surprise fees may be interesting but off-target unless cost-effectiveness is part of the conclusion.
To eliminate these, simulate the different outcomes and ask one question: does the conclusion genuinely rise or fall based on those outcomes? If not, it is likely a decoy.
4. If You Are Really Stuck: Invert
On hard questions, when binary checks and extreme tests are not enough, invert your perspective.
Take a candidate answer and ask: if I were writing a stimulus where this answer was clearly correct, would that stimulus look like the one in front of me? If the fit is clean, that is a strong signal. If the answer feels only loosely connected and seems better suited to a different argument, it is probably wrong.
This is especially useful when two answers both feel plausible. The right one usually feels structurally inevitable, while the close-but-wrong one has a subtle mismatch in scope or target.
Practical note: This method works well for high-intuition test takers in timed sections. Even a 60-40 lean can be valuable when the better-fitting answer cleanly aligns with the argument's exact logic.
5. Apply Evaluate Skills to Other Question Types
Evaluate is the bridge between flaw and strengthen/weaken work. If a student struggles to prephrase strengthen or weaken answers, I often back them up to Evaluate questions first, because this forces them to identify which missing facts would actually move validity up or down.
The overlap is direct. A common Evaluate answer asks whether an alternative cause exists. A common strengthener rules out competing causes. A common weakener introduces a rival cause. The same transfer happens with sampling quality, term definition, and analogy reliability.
The better you get at evaluating arguments in the abstract, the easier it becomes to retrieve the same moves quickly across other Logical Reasoning question types.