Executive Summary
As HealthTasks expands AI-powered educational analytics, institutions ask a practical question: which frontier model produces the best insights? Rather than relying on vendor leaderboards, we built a repeatable evaluation harness that measures insight quality and the operational efficiency required to produce it.
Across competency evaluations, skills checkoffs, clinical performance, curriculum analytics, and live Agents prompts, one conclusion held: there is no single winner. Gemini 3.6 Flash and Grok 4.5 excel in different dimensions—so HealthTasks supports both.
Lower combined Agents cost
Grok 4.5 vs Gemini 3.6 Flash on paired prompts
Fewer tool / LLM steps
Consistent across analytics and conversational runs
Complementary model profiles
Presentation polish vs statistical rigor — both supported
Inference cost by benchmark run
Same Agents pipeline and datasets — lower is better
Methodology
Every analytics request executed through the same HealthTasks Agents pipeline. The same datasets were analyzed independently by both models before comparing outputs. For each run we measured tool and API calls, input and output tokens, latency, inference cost, statistical reasoning quality, executive-summary quality, presentation readability, and identification of caveats or data limitations.
Evaluation spanned three vectors:
Insight quality
Statistical reasoning, executive summaries, presentation clarity, and identification of caveats or data limitations.
Operational efficiency
Tool and API call counts, input/output tokens, wall-clock latency, and total inference cost under the same pipeline.
Interpretive stance
Whether the model emphasizes polished ranking and narrative, or sample size, confidence, bias, and conservative recommendations.
Qualitative Profiles
Gemini 3.6 Flash
Gemini consistently produced polished reports that were easy to scan—executive summaries, well-structured narratives, clean organization, and attractive Canvas layouts. Outputs were immediately approachable for administrators and faculty reviewing institutional data.
Grok 4.5
Grok behaved more like a data analyst. Rather than simply ranking metrics, it frequently highlighted sample-size limitations, confidence in conclusions, data completeness, AI-generated versus educator-generated evaluations, and potential biases from incomplete datasets. Recommendations tended to be more conservative and statistically grounded.
The models rarely disagreed on the underlying data. They differed in how they interpreted it—Gemini emphasizing strongest performers and presentation; Grok emphasizing statistical confidence, sample sizes, and appropriate interpretation. For educational analytics, both perspectives provide value.
Efficiency Results
Across multiple benchmark runs, Grok demonstrated a significant efficiency advantage—fewer tool calls, lower token spend, and lower inference cost—while wall-clock latency stayed comparable.
Analytics benchmark #1
| Metric | Grok 4.5 | Gemini 3.6 Flash |
|---|---|---|
API Calls Grok used 50% fewer tool calls | 5 | 10 |
Input Tokens Comparable input volume | 357,945 | 380,136 |
Output Tokens Grok produced roughly half the output | 1,759 | 3,478 |
Total Cost Grok ~30% lower cost | $0.4159 | $0.5963 |
Execution Time Gemini ~4.5s faster | 40.68 s | 36.16 s |
Analytics benchmark #2
| Metric | Grok 4.5 | Gemini 3.6 Flash |
|---|---|---|
API Calls Grok used 56% fewer API calls | 4 | 9 |
Input Tokens Grok roughly half the input tokens | 86,343 | 166,758 |
Combined Tokens Nearly 2× fewer total tokens for Grok | 88,248 | 168,484 |
Total Cost Grok 47.8% lower inference cost | $0.1373 | $0.2631 |
Execution Time Only 1.8s additional latency for Grok | 31.69 s | 29.93 s |
Conversational Agents — Prompt A
Hours-versus-evaluations correlation with scatter visualization in the live Agents product.
| Metric | Grok 4.5 | Gemini 3.6 Flash |
|---|---|---|
LLM Steps Grok used fewer steps | 2 | 3 |
Input Tokens One Gemini step alone ingested ~181K tokens | 26,519 | 208,939 |
Output Tokens Comparable output length | 918 | 1,410 |
Total Cost Grok ~5.6× cheaper on this prompt | $0.0581 | $0.3240 |
Execution Time Nearly identical wall-clock | 22.2 s | 22.7 s |
Conversational Agents — Prompt B
Two-week clinical faculty priority plan.
| Metric | Grok 4.5 | Gemini 3.6 Flash |
|---|---|---|
LLM Steps Grok used half the LLM steps | 2 | 4 |
Input Tokens Grok ~58% fewer input tokens | 31,586 | 75,502 |
Output Tokens Grok produced slightly more structured plan text | 1,862 | 1,540 |
Total Cost Grok ~41% lower cost | $0.0739 | $0.1248 |
Execution Time Gemini faster on this faculty prompt | 38.6 s | 30.5 s |
Agents pair totals
- Combined cost: $0.132 (Grok) vs $0.449 (Gemini) — roughly 3.4× lower for Grok
- Combined time: ~61s (Grok) vs ~53s (Gemini) — comparable wall-clock
- Both models agreed on core classroom signals (for example ICU near-perfect hours and evaluations; NUR125 weak on key competencies around 13%)
- Grok answered the correlation question more tightly at classroom grain with a lean scatter Canvas; Gemini produced a richer exception taxonomy and competency-themed coaching plan
Discussion: Coding Benchmarks vs. Educational Analytics
On general-purpose software engineering leaderboards such as Datacurve's DeepSWE benchmark, both models sit in a competitive mid-frontier band. On DeepSWE v1.1 (mini-swe-agent), Grok 4.5 [high] scores 54%±2% Pass@1 and Gemini 3.6 Flash [high] scores 49%±5% Pass@1. DeepSWE measures long-horizon code refactoring across multi-file repositories—rewarding heavy reasoning and logical endurance on contamination-free software engineering tasks.
Educational analytics agent workflows demand a different profile: disciplined tool use, grounded statistical interpretation, schema-faithful Canvas synthesis, and cost-efficient token packing on institutional datasets. In that setting, the models diverge less on “who is smarter” and more on interpretive stance and operational economics—Gemini toward polished executive presentation, Grok toward conservative statistical framing and lower run cost—while often agreeing on the same underlying classroom signals.
DeepSWE ranking alone therefore does not dictate the right default for agentic educational analytics. Institutions should match the model to the job: presentation density, analytical caution, latency tolerance, and inference budget.
Why HealthTasks Supports Both
Rather than forcing institutions into a single AI provider, HealthTasks gives organizations flexibility to choose the model that fits their workflows, governance policies, and preferences.
Choose Gemini when you want
- Fast, polished executive summaries
- Administrator-friendly report presentation
- Dashboard-style Canvas layouts
Choose Grok when you want
- Statistical rigor and conservative interpretation
- Explicit caveats around sample size and bias
- Lower inference cost and fewer tool calls
Conclusion
HealthTasks Agents already supports multiple foundation models behind a common analytics pipeline. As new models emerge, we will continue benchmarking them with the same methodology—measuring not only how well they write, but how well they reason, how efficiently they operate, and how effectively they help educators make better decisions.
Choice is a feature. Different institutions have different priorities, and HealthTasks is committed to giving educators the flexibility to select the AI that works best for them.