Seven Ways AI Can Get an FE Problem Wrong
Published studies show real engineering problem-solving ability, but also reveal failure patterns that make an unverified chatbot a risky answer key.
A plausible solution can still be wrong
The wrong first question about AI and the FE exam is, "Can it pass?"
Published results make that question tempting. One 2023 paper reported GPT-4 at 70.9% on an FE-oriented evaluation. Another study using FE Mechanical practice questions reported 76% for GPT-4 and 51% for GPT-3.5. Those numbers are not a current universal ranking, an NCEES result, or proof that a chatbot is a safe answer key.
Models and test sets change. Some evaluations convert visual material into text or exclude what a system cannot read. A single percentage also hides answers that look rigorous but teach the wrong model.
I reviewed the literature with a practical question: how does a plausible solution fail?
1. It can misread the physical system
Engineering starts before the equation. A solver must identify the object, boundary, loading, direction, topology, and state.
Engineering-physics research found failures to construct an accurate physical model. Circuit-analysis work similarly identified recognition errors such as incorrect source polarity. Clean algebra cannot repair a wrong model.
Student check: Before asking for a solution, write the system boundary and governing phenomenon yourself. Compare the AI's interpretation with the diagram, not with your hope that it understood the diagram.
2. It can lose information in a diagram
The FE is not purely textual. Figures can encode dimensions, directions, supports, elevations, labels, or connectivity. The FE Mechanical study explicitly identified text-only input as a major limitation in its setting. Multimodal models have improved since that study, but an image-capable model can still overlook a small arrow, swap a label, or infer a connection that is not present.
Student check: Transcribe relevant labels into a givens list. If the model cannot reproduce the geometry and directions, stop before calculating.
3. It can invent a reasonable-looking assumption
Real questions sometimes require an implicit convention. Badly copied questions can also be incomplete. An AI system is often optimized to continue with an answer rather than pause at uncertainty.
The engineering-physics study reported bad assumptions about missing data. Research on unreasonable math problems found that models may detect a defect yet still hallucinate.
Student check: Require a separate "assumptions" block. Ask which assumption is supplied, which is conventional, and which was introduced only to make the problem solvable. An unsupported assumption is not a harmless detail.
4. It can select a valid formula outside its valid conditions
Many FE traps are not about remembering an equation. They are about knowing when it applies: steady versus unsteady, ideal versus nonideal, small deformation versus a different regime, laminar versus turbulent, or a particular support and loading case.
Language models are good at retrieving familiar equation patterns. That strength can become a weakness when surface wording points to a familiar formula but the conditions do not.
Student check: Ask for the governing principle first, then list every condition required by the chosen relationship. Verify the relationship against the current FE Reference Handbook or a trusted text. Do not accept a formula merely because its symbols match the givens.
5. It can mishandle signs, units, and reference frames
An answer may use a correct equation and still fail through a reversed direction, mixed unit system, absolute-versus-gauge pressure mistake, degree-versus-radian setting, or inconsistent datum.
These errors are especially easy to miss in fluent prose because each line appears locally reasonable. They also compound: one early sign convention can alter several later steps.
Student check: Keep units on every line, state the positive direction, and perform an independent dimensional check. Draw reference frames yourself and estimate sign and magnitude before calculating.
6. It can make a small calculation error inside a strong explanation
The engineering-physics analysis found calculation errors. A detailed response can create false trust: a model may derive the right setup, substitute incorrectly, and confidently explain the result.
Conversely, a model can occasionally reach the right choice through flawed reasoning. That answer is dangerous for study because you remember the path, not only the letter.
Student check: Recalculate independently with your approved calculator. Verify intermediate values, not just the final option. If a second prompt produces a different result, treat that disagreement as a warning, not a vote to keep querying until one answer feels right.
7. It can sound certain when the evidence is weak
The FE Mechanical paper noted inconsistency and confidently incorrect answers. Hallucination research more broadly describes outputs that are coherent but factually or logically wrong. Confidence in the writing is not calibrated confidence in the engineering.
AI can also invent a handbook section, page, citation, or rule. Even a real citation may not support the claim attached to it.
Student check: Open every cited source. Search the current handbook yourself. Ask the model to identify uncertainty and an alternative solution path, but do not treat self-critique as independent verification; it is still the same system generating more text.
The lesson is not "never use AI"
A 2025 randomized undergraduate physics study found that a carefully designed AI tutor produced greater learning gains in less time than comparison active-learning lessons. The intervention used expert-developed content, structured prompts, scaffolding, and pedagogical design, not an unrestricted chatbot.
That distinction is the product lesson I care about. AI is useful when it helps a learner retrieve prior knowledge, receive a graduated hint, explain a misconception, or build a focused plan. It becomes risky when fluency is allowed to impersonate verification.
My preferred FE workflow is:
An AI tutor should make this loop easier. It should not remove the loop.
- Attempt the problem without AI.
- Record the governing principle and givens.
- Ask for a hint before a full solution.
- Compare the proposed method with the current handbook.
- Recalculate independently.
- Check units, signs, limits, and answer plausibility.
- Retest the same concept later with a genuinely different, original problem.
Primary sources
Build the plan from your own baseline
A diagnostic cannot predict your result, but it can show which topic is the best next use of your study time.
Take the free diagnostic