A friendly interface can conceal wrong answers; critical review and willingness to rebuild are needed.
Proposed mechanism: Test answers against purpose and participant needs rather than interface preference.
Limit: Internal example without published error rate.