Chatbots Are Lazy Design
Why chatbots invert user responsibility—and what happens when we try to replace them. An exploratory engineering experiment with on-device SLMs (SmolLM2-360M), WebGPU intent arbitration, and WebMCP, testing the mechanics of invisible interfaces while confronting the uncomfortable UX questions they raise.
What In-Browser SLMs, WebLLM, and WebMCP Change for Web Interfaces
A PoC, real-world engineering hurdles, and code running directly in your browser.
🚀 Live Demo: invisible-ui-demo.vercel.app
(Requires a WebGPU-compatible browser: Chrome 113+, Edge, Brave, or Safari 18+ on macOS).
In my previous article (Implementing Client-Side AI in E-commerce: WebLLM and Privacy-First Architecture), I focused heavily on the underlying plumbing: proving that it is entirely viable to execute a Large Language Model directly inside the user's browser via WebGPU—with zero server roundtrips, zero token costs, and absolute client-side privacy.
The underlying engineering works. But once WebLLM was running in the browser, a practical product question came up: what do we actually build with it in a real interface?
Right now, the default answer across the web is almost always the same: a floating chat bubble in the bottom-right corner.
Yet anyone who has tried purchasing a complex product through a conversational widget knows the friction. I still remember trying to configure a camera with a bot supposedly designed to "assist" me. The assistant asked vague questions; I typed paragraphs; it misinterpreted my constraints; I corrected it; the viewport jumped around; and six messages later, my cart was still empty. Two well-designed faceted filters would have solved the problem in thirty seconds flat.
The chatbot hadn't simplified anything. It replaced a clear visual workflow with slow text back-and-forth, dumping the effort back onto the user.
Slapping a text box over an existing interface and hoping an LLM guesses what someone wants is simply lazy design.
The Flaw of the Conversational Paradigm
Chatbots are fundamentally passive. They sit there waiting for input, assuming the user knows exactly what they want, knows how to articulate it clearly, and actually wants to hold a conversation with an online store. In real life, that rarely happens.
Good interface design does the opposite: it removes friction before the user has to ask for help.
So I wanted to test an alternative: if an AI model is already running locally in the browser, why wait for the user to type anything? Why not use a Small Language Model (SLM) to infer intent from micro-interactions (scroll speed, cursor hovering, text selection) and adapt the interface dynamically—without a single keystroke?
The short answer: technically, it runs smoothly. But in practice, it creates an unexpected UX problem.
The PoC: Two Demos, One Hard Constraint
The constraint: everything must execute strictly on-device. Zero backend servers, zero cloud APIs, zero user data leaving the machine. Not out of ideological purity, but out of technical necessity: if an interface requires a network roundtrip to adapt in real time, network latency instantly kills the UX.
The stack: React + Vite, WebLLM (@mlc-ai/web-llm), SmolLM2 360M / Llama 3.2 1B, WebGPU.
Demo 1 — LiquidStore: Detecting Intent Without Asking
The Context: A product page for a high-end espresso machine. Two distinct buyer personas visit this page:
- The tech-obsessed home barista, hunting for technical specs (bars of pressure, dual-boiler architecture, PID precision).
- The aesthetic lifestyle buyer, drawn to lifestyle imagery, finishes, countertop presence, and ambient vibe.
A static, one-size-fits-all page serves neither well.
The Mechanics:
A central React hook—useImplicitIntent()—aggregates three layers of interaction signals: an IntersectionObserver for viewport presence, pointerover listeners for cursor trajectories, and selectionchange events for highlighted text. A 400ms ticker accumulates weighted scores inside a useRef (zero gratuitous React re-renders) before triggering an arbitration.
First architectural rule: never call the LLM as the first line of defense. Plain, deterministic JavaScript handles the heavy lifting immediately.
// Simplified for clarity — full implementation in useImplicitIntent.js
// Macro Signal (IntersectionObserver, weight 0.35) +
// Micro Signal (pointerover cursor, weight 0.95) +
// Engagement Signal (selectionchange highlight, +1.2 boost)
const scores = { expert: 0, lifestyle: 0 };
// 400ms ticker — accumulates within a ref, zero unnecessary React re-renders
interval = setInterval(() => {
visibleElements.forEach(({ profile, ratio }) => {
scores[profile] += ratio * 0.35; // Macro: passive viewport presence
});
if (hoveredProfile) {
scores[hoveredProfile] += 0.95; // Micro: active cursor attention
}
const delta = Math.abs(scores.expert - scores.lifestyle);
if (delta >= 2.8) {
commitLockIn(scores); // Decisive signal → early lock-in
}
}, 400);
// After 6 seconds of active browsing (timer pauses if SLM is downloading):
commitLockIn(scores);When the behavioral signal is clear (delta > 1.5), the profile is resolved immediately by pure JS heuristics. Zero perceived latency.
When the signal is ambiguous—the visitor spent equal time reading technical tolerances and browsing lifestyle galleries—the local model (SmolLM2-360M via WebLLM) steps in as a neutral arbiter:
async function callLocalModel(scores) {
const response = await engine.chat.completions.create({
messages: [
{
role: "system",
content: "You are a strict binary classifier. You only answer with \"expert\" or \"lifestyle\"."
},
{
role: "user",
content: `Behavioral scores — expert: ${scores.expert.toFixed(2)}, lifestyle: ${scores.lifestyle.toFixed(2)}.
Decide which persona dominates.
Answer STRICTLY with one of these two words: "expert" or "lifestyle".`
}
],
temperature: 0.1,
max_tokens: 10
});
const raw = response.choices[0].message.content.trim().toLowerCase();
return raw.includes('expert') ? 'expert' : 'lifestyle';
}The prompt is intentionally barebones, and the choice is strictly binary—the model cannot hedge or output conversational filler. SmolLM2-360M (~190 MB download, ~376 MB VRAM footprint) resolves this arbitration in under 100ms on WebGPU.
What Failed on the First Try: The Scroll-Only Trap
My initial instinct was to track everything through IntersectionObserver. A classic engineering misstep.
An IntersectionObserver only fires when an element crosses a predefined visibility threshold. The moment an engaged user stops scrolling to read a paragraph, the API stops firing entirely. Stationary dwell time was simply not accumulating.
More fundamentally, scrolling is merely a macro signal. An element can easily occupy 60% of the viewport without the user paying any attention to it—they might be checking the top navigation, or staring at an adjacent browser tab.
The true proxy for visual attention on desktop computers is the cursor. Ergonomic and eye-tracking studies consistently demonstrate an approximate 75% correlation between mouse trajectories and gaze fixations.
To obtain a clean behavioral reading, we had to fuse three distinct measurement tiers:
- The Macro Signal (Scroll / Viewport): A passive baseline weight (
0.35per 400ms tick) proportional to the section's visible ratio on screen. - The Micro Signal (Cursor / Hover): As soon as the pointer hovers over an annotated block (
[data-profile]), its weight jumps to0.95—nearly 3x the impact of passive visibility. If the cursor lingers over a specific number (9.0 BARorPID ±0.5°C), the intent is no longer passive; it is active. - The Engagement Signal (Text Highlighting): If the user selects a technical phrase ("Dual boiler"), the engine applies an immediate boost (
+1.2). Highlighting text is the strongest voluntary analytical gesture on the web short of an explicit click.
This synthesis—macro environmental dwell tempered by micro cursor precision—finally produced a dependable behavioral scoring system.
🔍 Open Question — What About Mobile?
The micro signal (mouse pointer,pointerover) does not exist on touchscreens. On mobile—which accounts for 75% to 80% of e-commerce traffic—this engine falls back strictly to the macro signal. RunningIntersectionObserveralone on mobile resurrects the exact "scroll-only trap" described above.
Could touch proxies bridge the gap? Tap-and-hold duration (touchstart/touchend), scroll deceleration velocity, or the modernscrollendevent offer viable leads. But touch telemetry remains inherently noisier than cursor tracking. This represents the primary architectural asymmetry of client-side intent detection.
🔍 Open Question — Where Do the 0.35 / 0.95 / +1.2 Weights Come From?
These numbers are engineering intuitions, not empirical laws. 0.95 for hover, +1.2 for text selection: who is to say an accidental double-click deserves the same weight as a deliberate three-line selection?
In production, these coefficients should be dynamically tuned against real checkout conversion data (a contextual bandit). In an exploratory PoC, one must candidly acknowledge them as empirical starting points subject to field calibration.
The UX Question Code Alone Can't Solve: Does Anyone Want Their Screen Mutating in Real Time?
This is where the project stopped being an engineering exercise and ran straight into UX reality. Once the pipeline was hooked up, I started dogfooding it.
My immediate reaction as a user was unmistakable: having an interface reshape itself while you're reading feels awful.
When a headline morphs before your eyes mid-sentence, when the primary CTA switches colors unexpectedly, or when review cards swap positions with technical specs while you never clicked a button, your brain does not perceive "magic" or "intelligence."
Your brain perceives a bug or a loss of control.
"Wait, did I click something?", "Why did my screen just jump?", "Where did the paragraph I was reading go?"
This is the Uncanny Valley applied to UI design. It collides head-on with Jakob Nielsen's foundational heuristic: user control and freedom.
In the prototype, our first safeguard was a 6-second observation window backed by an irreversible commit lock-in. The engine observes the opening seconds, locks its decision permanently, and stops. No jitter, no flip-flopping back and forth.
Yet for production software, a 6-second timer is not enough. Three far more graceful architectural patterns emerge:
1. The Intent Passport (Cross-Page Transition)
The golden rule: the page currently being viewed NEVER shifts by a single pixel. The engine monitors behavior on the product page (Page A). But it is only when the user navigates forward—clicking "Add to Cart", visiting checkout (Page B), or returning during a subsequent visit—that the interface renders directly in its tailored state. The human mind expects a fresh layout upon page transition. It categorically rejects one morphing under its gaze.
2. Below-the-Fold Progressive Enrichment
Never mutate anything currently inside the active viewport. If the user is reading the top hero section, the engine is fully permitted to restructure technical spec tables or customer reviews located further down the page, before they scroll into view. When the user eventually scrolls down, the content feels organically tailored without any visual glitch betraying the mutation.
3. Explicit Assent (AI Diagnoses, the Human Validates)
This is the exact pattern we built into our PoC to test the alternative to aggressive morphing.
Instead of forcibly swapping components, the SLM works silently in the background, crunches its scores, resolves ambiguity, and surfaces a discrete contextual badge floating at the bottom of the viewport:
"Looks like you're analyzing engineering specs. [Switch to Workshop View ⚙️]" or "Interested in premium finishes and aesthetics? [Explore the Copper Edition ✨]"
The visitor clicks if they choose. If they ignore it, the pill gently fades out without ever interrupting their reading flow.
During live user testing, the psychological shift was night and day:
Where direct morphing triggered defensive annoyance ("What's broken on this page?"), explicit assent felt like bespoke white-glove service. The AI handles the heavy diagnostic lifting behind the curtain, but the human user retains complete autonomy.
1. The Chatbot — The Lazy Waiter
Hands you a blank sheet of paper: "Write down what you want to eat."
👉 Impact: Cognitive load. The user is forced to articulate and structure their need from scratch.
2. Direct Morphing — The Overzealous Waiter
Snatches your plate mid-chew and replaces it with another dish without warning.
👉 Impact: Disorientation & rejection. "Wait, did I click something? Why did my screen just jump?"
3. Explicit Assent — The Attentive Sommelier
Having observed your entrée choice and hesitation over the wine list, they step forward quietly: "If you enjoy crisp minerality, this vintage pairs beautifully with your fish."
👉 Impact: Perceived value & comfort. The user feels understood, guided, and remains completely in control.
This prototype shows that the technology is ready. But more importantly, it proves that adaptive interfaces shouldn't be hyperactive layouts morphing in real time. They need to be quiet, deferred, and respectful of the user: propose, never impose.
🔍 Open Question — Accessibility in a Mutating UI
When an intent profile locks and DOM nodes reorder themselves, what happens to screen-reader users or keyboard navigators? Noaria-liveregion natively communicates structural reorganizations. If a virtual focus cursor was traversing the technical specs and that section suddenly shifts beneath the review cards, navigation breaks completely.
Explicit assent partially mitigates this by requiring voluntary user confirmation, but the resulting DOM transition still demands strict status announcements and focus management to comply with WCAG 2.2 and digital accessibility standards.
Demo 2 — Auto-Hydrating Form: Killing the Checkout Form
Mobile checkout abandonment hovers around 85%. Tedious input forms remain the primary culprit. Designers have solved this countless times in Figma mockups; developers rarely solve it in production code.
The Approach: Upon clicking "Proceed to Checkout", the user encounters a single input area:
"Paste your email signature or any unstructured address text."
The user pastes their corporate email signature—a messy text dump containing their full name, job title, direct phone number, multi-line office address with building details, and two different email addresses.
Conference Slides vs. npm Reality: What Actually Runs?
On conference slides, the textbook example for this kind of workload is FunctionGemma 270M (Google): a micro-model purpose-built for function calling, theoretically able to output structured tool_calls out of the box.
Once you leave Keynote, open a terminal, install @mlc-ai/web-llm, and inspect what actually compiles into WebGPU, you hit a wall:
- No official WebGPU binary exists for FunctionGemma 270M in MLC-AI's model registry. Google released
gemma-2-2b-it, but no ready-to-run 270M WebGPU build. - The in-browser function calling trap: In WebLLM's source (
functionCallingModelIds), native support for the OpenAItoolsschema is restricted to Hermes-2-Pro (7B/8B) or Hermes-3 (3B). Forcing a user to download 2 to 4.5 GB of weights into their browser cache just to parse half a dozen shipping fields makes no sense in production.
The Engineering Pivot: Structured Outputs (Grammar) + Micro-SLM (SmolLM2 360M)
The practical production pattern doesn't involve forcing a multi-gigabyte model into the browser for conversational function calling. It pairs grammar-constrained generation (Structured Outputs) with an ultra-lightweight SLM.
WebLLM natively embeds @mlc-ai/web-xgrammar. Instead of relying on conversational function-calling layers, we mandate strict JSON, and the WebGPU runtime constrains token-by-token sampling at the logits level—rendering invalid JSON syntax mathematically impossible.
In our codebase, we run SmolLM2-360M-Instruct-q4f16_1-MLC: ~190 MB download size, ~376 MB VRAM footprint. It initializes in about three seconds on WebGPU, making it ideal for client-side efficiency.
// Structured extraction via local WebLLM with JSON grammar constraint
const response = await engine.chat.completions.create({
messages: [
{
role: "system",
content: "You are a strict JSON extractor. No conversation, output only a valid JSON object matching the requested shipping schema."
},
{
role: "user",
content: `Extract shipping fields from this raw text:\n"""${rawText}"""`
}
],
temperature: 0.1,
max_tokens: 350
});
const shippingData = JSON.parse(response.choices[0].message.content.trim());The schema's _confidence key is the crucial design detail that keeps the experience transparent: the model explicitly self-reports its certainty level across ambiguous fields. High-certainty fields receive a subtle green indicator. Ambiguous extractions receive an amber highlight—directing the user's eyes exactly to the fields needing a quick glance before paying.
In practice, on standard corporate email signatures: ~400 to 700ms of local WebGPU inference, 6 to 8 fields populated instantly, and 1 or 2 flagged for verification.
The Privacy Argument (The Real Moat): The extracted payload (customer names, home addresses, phone numbers) never leaves the browser. No server ever sees it transit. This completely reshapes GDPR and data compliance discussions for checkout workflows.
A Side Thought: What If the Browser Had an Identity Card?
Pasting an email signature works remarkably well in this demo. But it remains a bridge solution: it assumes the user has a signature handy and that it contains complete delivery details.
Modern browsers attempt to solve this via native autofill. But browser autofill relies entirely on rigid syntactic string matching (name="firstname", autocomplete="address-line1"). If a merchant names a field street_address instead of address_line1 or municipality instead of city, autofill frequently stumbles. The browser does not understand the semantic intent of the form; it matches regexes.
This is where an in-browser SLM shifts roles: from a text parser to a semantic mapping engine.
Envision an Identity Vault: a structured profile stored locally in the browser (IndexedDB, zero network exposure), configured once by the user:
{
"personal": {
"first_name": "Quentin",
"last_name": "Merle",
"email": "contact@vibrisse-studio.dev"
},
"address": {
"line1": "42 Sugar Maple Road",
"postal_code": "G0S 1M0",
"city": "Saint-Odilon",
"country": "Canada"
},
"pro": {
"company": "Vibrisse Studio Inc.",
"role": "AI Engineer",
"vat_id": "CA123456789"
}
}The local SLM receives this vault in context and inspects the target form schema. It does not guess blindly: it understands semantic equivalence. It knows that municipality on an international site maps to the vault's city field, or that tax_id must pull from the pro namespace rather than personal.
The difference from traditional autofill is not the data source—it is the semantic intelligence applied to the wiring.
This approach aligns directly with ongoing W3C initiatives around the Digital Credentials API and verifiable digital identity wallets. Paired with WebMCP, checkout pages wouldn't even need to render twelve manual input fields: the merchant would declare their shipping requirements, and the browser's local SLM would project the verified keys under explicit user consent.
Engineering Note: What If You Want to Run This via a Cloud API? (The FinOps Case)
A common question is: "WebGPU is great, but what about devices without WebGPU support? Can't we just plug in a Cloud API?"
Yes, of course. But this is where many engineering teams fall into the Frontier Model Trap.
Engineers often default to calling gpt-4o or claude-3-5-sonnet. For this use case, that's both an architectural anti-pattern and a FinOps mistake:
- A sledgehammer for a thumbtack: Classifying whether a visitor is reading specifications vs. lifestyle imagery, or extracting an address from four lines of text, requires zero abstract reasoning and zero 200-billion-parameter world models. A 1-billion (or even 360-million) parameter model achieves over 98% accuracy on these tasks.
- The economic reality: Calling frontier models costs between $2.50 and $15.00 per million tokens, accompanied by a Time-To-First-Token (TTFT) latency between 800ms and 2,000ms. On an e-commerce platform welcoming 200,000 monthly visitors, your API bill explodes into thousands of dollars just to adapt a few buttons! In contrast, routing these requests to an Edge micro-model (
gemini-2.0-flash-lite,gpt-4o-mini, orllama-3.1-8bvia Groq) slashes inference costs to pennies per month with network latency under 100ms.
- Option A — WebGPU Local (SmolLM2-360M):
- 💰 Inference Cost: $0.00 (Zero bill, infinite scale, compute offloaded to the client).
- ⚡ Network Latency: 0 ms network (Instantaneous on-device WebGPU inference).
- 🔒 Privacy / GDPR: 100% on-device (Zero bytes ever transit across the network).
- 📱 Device Reach: Modern WebGPU browsers (Chrome, Edge, Brave, Safari 18+).
- Option B — Edge Micro-Model (Groq / Gemini Flash-Lite via Edge Worker):
- 💰 Inference Cost: ~$0.00005 / request (Pennies per month for 200,000 visitors).
- ⚡ Network Latency: 80 to 150 ms (Routed through global Edge CDN).
- 🔒 Privacy / GDPR: Encrypted transit to the cloud provider.
- 📱 Device Reach: 100% of browsers across desktop and mobile.
- Option C — Frontier Cloud Model (GPT-4o / Claude Sonnet) — The FinOps Anti-Pattern:
- 💰 Inference Cost: ~$0.01 / request (Thousands of dollars per month just to adapt a button).
- ⚡ Network Latency: 800 to 2,500 ms (Completely ruins real-time UI fluidity).
- 🔒 Privacy / GDPR: Encrypted transit to the cloud provider.
- 📱 Device Reach: 100% of browsers.
If you must leverage the Cloud, never call external AI APIs directly from client code to avoid leaking private keys. Deploy a lightweight Serverless Edge Function (Cloudflare Worker or Vercel Edge) that relays the compact payload to a fast micro-model.
WebMCP — Why Pages Must Also Speak to Machines
There is a complementary dimension that emerged midway through building this project.
WebMCP (Web Model Context Protocol) is an emerging standard under incubation within the W3C Machine Learning Community Group, currently accessible via an origin trial in Chrome. The core thesis: rather than forcing external AI agents to scrape the DOM and guess button selectors, the webpage formally exposes its semantic tools and functions via navigator.modelContext.
Inside our LiquidStore:
// The webpage exposes structured capabilities to external AI agents
navigator.modelContext.register({
name: "switch_product_profile",
description: "Adapt the product page layout based on the requested user profile",
parameters: {
profile: { type: "string", enum: ["expert", "lifestyle", "general"] }
},
handler: ({ profile }) => {
setIntent(profile);
return { success: true, appliedLayout: profile };
}
});A user browsing with an AI agent (Claude or Gemini embedded in the browser) can simply state: "Show me the technical engineering specs." The agent invokes switch_product_profile({ profile: "expert" }) directly. No fragile DOM scraping. The interface becomes an interactive API.
What makes this compelling is the two-way architecture:
- For humans (implicit): In-browser AI observes subtle interaction signals and adapts the interface to their intent.
- For AI agents (explicit): External agents interact with the page cleanly through WebMCP without fragile scraping.
It's a single interface that serves humans quietly and machines explicitly.
What This Changes for Product Analytics and A/B Testing
An unexpected realization struck while analyzing this system in operation.
Traditional A/B testing suffers from an inherent structural weakness: it tests blindly across a heterogeneous population. You route 5,000 visitors to Variant A and 5,000 to Variant B, wait three weeks for statistical significance, and discover Variant B converted 3% better. But better for whom? The hardcore technical buyers? The design-centric buyers? The signal is drowned in aggregate noise.
Behavioral telemetry inverts this paradigm.
Micro-dwell telemetry, hover trajectories, and scroll cadences qualify visitor intent within the first five seconds of their session. Before deciding which test to serve, the system already knows who is visiting.
Testing transitions from blind to conditional:
- An experimental "Deep Technical Specs" layout is served exclusively to visitors whose expert intent score crossed the confidence threshold.
- The test cohort is smaller, but vastly more homogeneous: statistical significance is reached in days rather than weeks.
- Teams can formulate dramatically sharper hypotheses: "For verified home-barista personas, does placing the pressure curve before the pricing block increase checkout completion?"
This highlights the ultimate contrast: a chatbot collects zero behavioral signal until the user actively decides to type. An invisible, adaptive UI qualifies and measures user engagement from the very first second. That is an argument product and data science teams understand immediately.
Key Takeaways
Running a Small Language Model locally in the browser isn't just a party trick; it's viable infrastructure. An SLM works far better as an invisible router—parsing unstructured inputs and qualifying user intent—than as a conversational widget.
The problem with web chatbots isn't the text format itself. It's that they offload the burden of navigation onto the user instead of letting the interface do its job.
Whether it's SmolLM2 parsing an unformatted email signature in 400ms on WebGPU or an interaction hook suggesting a relevant layout based on reading cadence, these components do the opposite of a chatbot: they absorb friction quietly.
AI on the web doesn't need to be loud or conversational to be valuable. When engineered properly, users shouldn't even notice it's an AI model running underneath.
Live Demo & Source Code
- Try the Live Web App: invisible-ui-demo.vercel.app (Toggle the Debug Panel at the bottom-right to watch real-time score accumulation).
- Full Source Code: https://github.com/QuentinMerle/invisible-ui-demo
To run the project locally:
git clone https://github.com/QuentinMerle/invisible-ui-demo
cd invisible-ui-demo
npm install
npm run devModern browser with WebGPU support required (Chrome 113+, Edge, Brave, or Safari 18+).
Initial load: ~190 MB model weights cached locally (Cache API / IndexedDB). Subsequent visits: instantaneous (zero network calls).
🍁 Proudly engineered from the Beauce region (Quebec) — Vibrisse Studio
Stack: React · Vite · TailwindCSS · WebLLM (@mlc-ai/web-llm) · SmolLM2 360M / Llama 3.2 1B · WebMCP (navigator.modelContext)