All work

Jamovi / Usability & Agent Experience (AX)

Jamovi is beloved for making statistics feel friendly. Across three studies, human think-aloud testing, expert PURE walkthroughs, and a 7-agent AI benchmark, I found that its calm surface silently lets users reach wrong conclusions.

0
phases
6 + 3 + 7
users + experts + agents
0
coded error events
0.00
SUS score
Jamovi One-Way ANOVA statistical analysis interface from the research sessions
Pufferfish mascotthe pufferfish trilogy

TL;DR

  • Low perceived task load (NASA-TLX 19.83) masked critical usability barriers (SUS 62.92): participants assumed their analysis was correct while generating invalid models.
  • Errors clustered in the ANOVA execution phase (64% of 100 errors); subtle icon animations without explicit text feedback allowed initial mistakes to cascade silently.
  • AI agents faced similar affordance gaps (only 1 of 7 completed autonomously), but adding two lines of contextual guidance increased screenshot interpretation accuracy from 28/50 to 42/50.

Context

Friendly design can hide sharp edges.

Jamovi is free, open-source statistics software positioned as the friendly answer to SPSS. Its pufferfish mascot, clean spreadsheets, and direct menus carry a mission: make statistics accessible to everyone, not just methodologists.

Approachability and statistical correctness make different promises. When a tool feels effortless, novices readily extend that trust to their output, assuming that a calm interface means a valid analysis.

The central research question: does Jamovi’s approachability come at the cost of diagnostic awareness? What happens when novices trust a silent interface to catch their methodological mistakes?

“It feels so intuitive to use, I thought I was doing everything right.”

Participant P03 · Think-aloud session (reported low subjective workload, yet generated misconfigured ANOVA output)

Role

Lead Researcher

Experimental design, moderation, qualitative coding

Scope

3-Phase Trilogy

Think-aloud protocols, PURE expert walkthrough, AI agent benchmark

Focus

Statistical UX

Investigating operational ease vs. cognitive diagnostic failure

Research questions

1

RQ1Does Jamovi’s friendly interface facilitate or hinder statistical accuracy?

Examining whether an approachable surface conceals critical methodological errors during One-Way ANOVA tasks.

2

RQ2Why does low subjective workload diverge from objective usability?

Explaining the paradox between low perceived effort (NASA-TLX 19.83) and marginal usability scores (SUS 62.92).

3

RQ3Can AI agents reliably operate the interface, and what does it reveal about design legibility?

Evaluating Computer Use agent navigation and the impact of minimal contextual cues on machine comprehension.

Methods

Three phases.

phase one

Usability Test & Think-Aloud

Moderated sessions with 6 STEM participants performing One-Way ANOVA tasks. 100 error events coded across 4 task phases.

n=6 STEMSUS 62.92NASA-TLX 19.83

phase two

PURE Expert Evaluation

3 UX experts independently evaluated 14 workflow steps (42 ratings) using PURE methodology alongside 10 Nielsen heuristics.

3 UX ExpertsPURE 19/42Heuristic Severity 1.67

phase three

Agent Experience Benchmark

7 AI agent models evaluated on autonomous task execution via Computer Use and screenshot statistical interpretation.

7 Agent ModelsComputer UseFDA Guidance Hint

Key findings

64%

of all errors concentrated in Task 3 (One-Way ANOVA execution), where variable configuration occurred.

Phase I Usabilityn=100 errorsTask 3 Peak
62.9 vs 19.8

SUS usability (62.92, marginal) sharply diverged from NASA-TLX workload (19.83, very low), showing novices felt confident while making critical errors.

Subjective-Objective GapLow Load, High Error
29/70

downstream errors propagated from ambiguous visual cues that flashed without explicit explanatory text.

PURE EvaluationSilent Error Cascade

Where the errors concentrate.

100 error events across moderated sessions were coded across the four task phases. Over 64% clustered in the ANOVA execution stage alone.

Error count across task phases T1 through T4, showing concentration in T3 ANOVA execution
Figure: Error Count by Task Phase (T1–T4, n=100 errors across 6 participants).
Task 1 (Preparation: 8 errors) & Task 2 (Descriptives: 14 errors): Clean spreadsheets and direct menus produced relatively low initial friction.
Task 3 (One-Way ANOVA Execution: 64 errors): 64% of all errors concentrated in variable mapping, where invalid drag-and-drop operations were silently accepted.
Task 4 (Assumption Checks: 14 errors): Ambiguous visual feedback led participants to trust statistically flawed results without verification.

Agent Experience (AX)

I tested seven AI models on the same ANOVA task. Only one completed autonomously.

In Phase III, I evaluated 7 AI models using Computer Use to drive Jamovi through the same One-Way ANOVA tasks. Six of seven failed to complete autonomously, encountering the exact same affordance traps that tripped human participants: silent drag mismatches, subtle icon cues, and unverified default selections. Only GPT-5.4-xHigh completed the full pipeline.

Next, I tested screenshot interpretation across models under default output versus with a two-line contextual hint (incorporating FDA equivalence testing standards). Accuracy surged from 28/50 to 42/50. Interface legibility and persistent context descriptions are low-cost interventions that benefit both humans and machines.

7 AI models evaluated via Computer Use · 1 completed autonomous analysis task

Jamovi One-Way ANOVA Welch table with FDA equivalence testing guidance notes
Figure: Contextual Guidance Intervention in Jamovi (Phase III). Adding two lines of statistical caveat into the results table raised AI interpretation accuracy from 28/50 to 42/50.
0/50

Default output

0/50

With two-line contextual hint

Screenshot interpretation accuracy across models: before vs. after textual guidance intervention

Recommendations ranked by empirical evidence.

Replace silent icon flickers with explicit error banners when variable types or distributions mismatch.

Evidence: Finding 3 (Silent Drag)

Introduce statistical linting to catch data entry errors before running hypothesis tests.

Evidence: Phase I (Data QA)

Implement progressive disclosure wizards for ANOVA workflows, surfacing assumption validation contextually.

Evidence: Finding 1 (Task 3 Peak)

Embed standardized statistical interpretation caveats directly in output panes to aid humans and AI agents.

Evidence: Phase III (AX Hint)

critical moderate polish

0.00

SUS

0.00

NASA-TLX (of 100)

0

PURE (of 42)

0

error events

0

of 7 agents completed

Limitations & Reflection

What I’d do differently.

Small samples bound every phase of this study: six users, three experts, seven agents. The patterns they reveal are consistent and mutually reinforcing, but individual metrics should be read in context rather than generalized as population statistics.

Error coding reliability was handled with an explicit codebook refined across multiple passes. Still, having independent multi-rater agreement validation remains a valuable methodological next step.

The agent benchmark is a snapshot of fast-moving model capability. The specific 1/7 completion number will evolve as models improve; the underlying finding will not: interfaces are sensory environments for machines too.

Finally, this research focused on diagnostic evaluation rather than a deployed production redesign. Every recommendation is evidence-backed and ranked by empirical findings; closing the loop with open-source Jamovi contributions or an intervention study is the natural continuation.

“The most dangerous usability problems are the ones that feel like nothing at all.”