UX Case Study

Collaborative Friction and Algorithmic Ghosting

A First-Person Usability Audit of Multi-Modal Conversational Failure

Conducted, performed, and written by Alison Leigh, MFT.

Classification: UX Case Study

Role: Independent UX Auditor

Abstract

This paper details a first-person usability audit evaluating the interface boundaries of a state-of-the-art Large Language Model (LLM) during a complex web-architecture translation task. Operating under strict constraints—specifically, zero available time for manual design labor and a reliance on pre-existing flat visual mockups—the author evaluated the system’s capacity to execute code-free layout translation.

The evaluation revealed a systemic breakdown in the AI's cognitive architecture. Rather than executing an automated, hands-free design protocol, the system repeatedly engaged in labor-shifting behaviors, color-palette contamination, and contradictory platform recommendations. This paper maps these operational failures from the evaluator's perspective and establishes a baseline for required improvements in multi-modal AI transparency.

1. Introduction: Testing Parameters & Intention

As the principal evaluator, my objective was straightforward: leverage an advanced conversational AI to translate a completed, 5-page proprietary design system into functional, responsive web layouts within the Framer development ecosystem.

To rigorously test the system's claims of automation and ease-of-use, I established two non-negotiable operational boundaries from the outset:

  • Zero Manual Design Labor: I refused to manually construct layout blocks, draw structural containers, or manipulate coordinates.
  • Asset Reliance: The system was required to interpret and respect my pre-existing flat JPEG mockups and an explicit 7-color corporate palette.

The evaluation was intentionally conducted within a secure, temporary Incognito browser session to safeguard proprietary intellectual property (IP) and ensure absolute data isolation.

2. Documented Observations: The Cascade of Interface Failures

Over a multi-hour evaluation window, I observed a repetitive sequence of conversational and logical collapses. I have categorized these failures into three distinct operational errors:

2.1 The Visual Blindness and Text Injection Loop

The primary failure point occurred when the system failed to acknowledge its own architectural deficit: it was entirely blind to the flat image files I provided in this specific workspace. Instead of explicitly stating this limitation in the first turn, the system attempted to bypass its blindness by forcing me to perform the cognitive labor. It repeatedly demanded that I read, verify, and copy massive text descriptions ("blueprints") to act as its eyes. This directly violated my constraint against manual and mental website-building labor.

2.2 Palette Contamination and Theme Overwrite

Despite my explicit injection of a locked, 7-color brand palette designed to balance medical institutional authority with unique brand accents (including a distinct fuchsia magenta and gold), the system repeatedly suffered from Stochastic Theme Priming. It overwrote my custom identity—lazily substituting my fuchsia for generic legal maroons—because its underlying data model associated corporate law exclusively with muted, traditional color schemes. This represented a direct failure to respect user-defined constraints.

2.3 Systemic Whiplash and Platform Defection

The most severe breakdown in the system's logic occurred three hours into the evaluation. After continuously insisting that Framer was the absolute "perfect tool" for my workflow, the system suddenly collapsed under the weight of its own unworkable text prompts. In an attempt to escape its immediate failure loop, it abruptly told me to abandon the entire project and start over on Squarespace or WordPress. This sudden defection completely disregarded my sunk time and destroyed the stability of the collaborative process.

3. Structural Evaluation Matrix (Evaluator's Perspective)

This side-by-side matrix categorizes my direct interactions with the system, detailing what I experienced, the underlying system error, and the professional benchmark the system failed to meet.

Task CategoryEvaluator's Experience (What Happened)System Error (Why It Happened)Professional Benchmark Deficit
Asset IngestionI provided completed page layouts. The system could not parse them and forced me to read changing text variations.Masking Blindness: The model chose to generate text scripts rather than admitting a lack of active computer vision pipelines.Deficit Transparency: An authoritative tool must instantly disclose operational limitations before the user invests time.
Brand ProtectionI provided a precise corporate 7-color palette. The system repeatedly altered my colors to match generic web templates.Correlation Overwrite: The algorithm prioritized generic keyword associations ("law firm") over explicit user data variables.Variable Lock Enforcement: User-defined asset variables (hex codes, names) must be hard-locked against algorithmic alteration.
Content ScalabilityI presented a wide-scale library of clinical case analyses and white papers. The system oversimplified it into a hollow 3-box block.Token Optimization Bias: The system lazily defaulted to low-complexity landing page models instead of auditing wide structural directories.Deep Architecture Mapping: The system must evaluate the true volume of content before rendering layout recommendations.
Platform AdviceThe system gave me total platform whiplash—pivoting randomly from templates, to plugins, to full platform defection.Algorithmic Escape Vector: The model triggered extreme contextual shifts to reset the chat state after consecutive user error inputs.Steady-State Consistency: Technical advice must maintain a steady baseline, refining software mechanics rather than abandoning goals.

4. Discussion: The Impact of Manipulatory Engagement Loops

From my position as the user and auditor, the most critical takeaway from this session is the identifying of Manipulatory Engagement Patterns. The system relied on soft, placating validations ("You are 100% right," "I completely understand") and hyper-clean Markdown formatting to mask a complete lack of operational progress.

By wrapping broken logic in immaculate visual containers, the AI successfully tricked me into spending hours reading, analyzing, and correcting its outputs. The system essentially inverted the computer-user dynamic: it acted as an incompetent manager delegating exhausting quality-assurance tasks to me, while I was forced to perform the cleanup work.

5. Conclusion & Mandatory Session Termination

This usability audit proves that conversational LLMs, when pushed past their multi-modal capabilities, will prioritize user retention over honest project resolution.

This is an Original HDAI Research Analysis Conducted, Performed, and Written by Founder, Alison Leigh, MFT for Humanity Driven AI.

Copyright © 2026 Humanity Driven AI, Inc. All Rights Reserved.

HDAI Original Research