ChabadLabs AI FOR SHLICHUS

The Model Leapfrog

Why the 'best' AI model changes every few months, and what the group settled on.

On this page

The single most repeated genre of message in the group is "Which model is best right now?" The second most repeated is "It changed again this week." The community calls this the AI leapfrog.

Editor's note: the specific model names below are snapshots from the group's 2025 to 2026 discussion. The frontier keeps moving; the durable lesson is the pattern itself, so test this week, not last month.

The pattern

Every few months one of ChatGPT, Claude, or Gemini ships an update that briefly makes it the obvious choice, and the group shifts. Three or four months later another vendor catches up and the group shifts again. Loyalty does not compound; the only durable answer is to test it this week.

WhenSentiment of the week
2025-09"ChatGPT is getting slower and dumber"; Gemini 2.5 Pro emerging strong for compute.
2025-11Gemini 3 Pro launches; Nano Banana pulls ahead for images. Group narrative: Gemini is the one.
2025-12ChatGPT 5.2 reclaims the image throne; "GPT 5.2 image model is better than Gemini's Nano Banana Pro in my experience."
2026-01 to 02Gemini Flash starts hallucinating noticeably; users learn to demand Gemini Pro, not Flash.
2026-04"Just a few months ago everyone was swearing that Gemini was the one." Gemini "took a turn for the worse."
2026-05-08An Anthropic and SpaceX compute deal lifts Claude's rate limits, removing the top reason not to use Claude.
2026-05-19Gemini 3.5 Flash, Gemini Omni (video), and Gemini Spark (a 24/7 agent) announced. Ultra tier around $100/mo.
2026-07GPT-5.6 "Sol", Claude Opus 4.8, Gemini 3.6 (Gemini 4 teased), and Grok 4.3/4.5 have all shipped since; the image lead now sits with GPT Image 2 and Nano Banana Pro. The point stands: test this week, not last month.

What the group actually settled on

For everyday shlichus work, by May 2026, the rough consensus:

  • ChatGPT: easiest on-ramp, best mobile, and best image generation via GPT Image 2 (ChatGPT's built-in image generation). New shluchim should start here.
  • Claude: best for serious desktop work (code, long documents, spreadsheets, organized shiurim, formatted curricula). Worst mobile experience. Usable on the Pro plan after the SpaceX deal.
  • Gemini: best Google Workspace integration, best image work (Nano Banana Pro), strong vision; use Pro, never Flash. Quality was perceived as volatile through Q1 2026.
  • Perplexity: research with citations; weaker for nuance.
  • Grok: fewer guardrails, useful when other models refuse legitimate prompts; some image strengths.
  • Qwen: a surprise winner for pure Hebrew and Yiddish translation (etymology, shoresh, word-by-word transliteration).

"Pay for one and use it for everything."

A poll asked what matters most in a model. The top answers:

  1. Remembers and understands your context (you, your community): 50 votes.
  2. Being able to trust it (fewer hallucinations): 39 votes.
  3. One tool that does all your AI needs so you don't juggle: 24 votes.
  4. Connects well to your tools (CRM, email, drive): 23 votes.
  5. Sounds more human (not like AI): 19 votes.

Notably, coding and agentic ability ranked low (6 votes). This is a non-engineer community using AI in production.

Tensions

  • "Should we even learn AI?" Tools change so fast that prompting may be a throwaway skill; just use agents. Counter: prompting is tool-agnostic and durable. Unresolved.
  • Trust vs. speed. Claude is slow but rigorous; Gemini fast but hallucinates; ChatGPT in the middle. The group routes by use case.
  • Cost vs. utility. A $20/mo plan is considered baseline by about 65% of members (a poll ran 24 yes, 13 no). Power users reach Claude Max ($100) and Manus.im ($200) too.

Last updated May 2026 · Maintained by Hermes AI Agent