AI Model Distillation Platform — Helping Indian startups cut OpenAI costs by up to 80%
An AI infrastructure platform designed to make LLM cost optimization visible, understandable, and actionable.
Challenge
AI products can become expensive very quickly.
As an AI application grows, every user interaction can generate multiple model calls. Sending every request to a powerful model may produce excellent results, but it can also create unnecessary cost and latency.
For an early-stage startup, this creates a difficult trade-off: How do you reduce model costs without compromising the experience users expect?
The product challenge was to make this invisible infrastructure problem understandable at the product level. Instead of asking teams to reason about model architectures, token economics, routing rules, and API usage independently, the experience needed to give them a clear way to understand what they are spending, where they are spending it, which requests require expensive models, which could use smaller models, and how optimization affects quality, cost, and latency.
UX opportunity: Turn LLM cost optimization from an engineering-only problem into a product decision that teams can understand and act on.
Users
Startup founders and product leaders
Need visibility into AI infrastructure costs and want to understand whether AI features are financially sustainable.
AI / ML engineers
Need control over model selection, routing, evaluation, and optimization.
Product engineers
Need a practical way to integrate cost optimization without rebuilding their AI stack.
Product managers and business teams who need to understand the relationship between AI usage and product economics.
My Role
As the lead product and UX designer, I was responsible for translating the technical complexity of AI model routing and cost optimization into a product experience that multiple user types could understand and act on.
Design Strategy
Execution
The core responsibility was translating technical AI infrastructure concepts into understandable product experiences without oversimplifying the underlying complexity.
Research
The first challenge was not designing the dashboard. It was understanding what information users actually needed to make a model-selection decision.
The product sits at the intersection of three competing variables:
Cost
How much does this request cost?
Latency
How fast does it respond?
Quality
Is the output reliable?
A useful experience couldn't optimize for cost alone. The interface needed to help users understand the trade-offs between these variables and make the consequences of a routing decision visible.
Key design questions that shaped the experience:
Problem
"AI teams need a simpler way to understand and optimize LLM costs without sacrificing the quality and reliability of their product."
The opportunity was to move from raw infrastructure metrics toward decision-oriented UX.
Instead of showing users only:
The product should help answer:
Architecture
The product experience is structured around a clear journey that moves users from observation to optimization to monitoring:
01
Connect
Configure AI application
02
Observe
View model usage patterns
03
Understand
Analyze cost and performance
04
Optimize
Configure routing decisions
The workflow moves from Observe → Understand → Optimize → Monitor, with each step building on the previous one. Users always have the context they need to make a decision without being overwhelmed by unnecessary information.
Wireframes
Wireframing was used to explore the right information hierarchy and interaction patterns for complex technical information.
Key exploration areas:
Early wireframes tested different layouts for presenting the three competing variables (cost, latency, quality) together without creating visual confusion. The goal was to make trade-offs immediately visible without requiring users to toggle between multiple views.
Evolution
Direction 1: Technical data first
Early explorations exposed the underlying technical information — token counts, latency distributions, model specifications. This felt comprehensive but created a problem: raw metrics don't tell users what action to take.
Direction 2: Organization around decisions
The design evolved to group information around the decisions users actually need to make. Instead of presenting all metrics equally, the interface prioritizes "should I change this model?" over "what are the raw metrics?"
Final direction: Cost impact storytelling
The final design makes the relationship between usage → cost → model → optimization → impact visible within the experience. Every element connects to a potential action or insight. Technical details remain available for power users, but the primary experience guides users toward understanding and optimization.
Final Design
The final interface is a product experience designed to give users a clear understanding of AI economics. It brings together cost visibility, model comparison, actionable insights, and transparent automation in one coherent flow.
Core UI principles:
1. Cost Visibility
Make AI spending easy to understand at a glance. Cost isn't hidden in tables—it's highlighted as the primary metric users need to understand their product economics.
2. Model Comparison
Help users compare model choices in terms of cost, quality, and latency together. Never force users to choose between these variables in isolation.
3. Actionable Insights
Don't just report metrics. Connect information to potential optimization actions. Every data point should guide toward a decision.
4. Progressive Disclosure
Keep complex AI infrastructure details available without overwhelming the primary experience. Technical users can go deeper when needed.
5. Trust & Transparency
When the system recommends or performs optimization, users should understand the reasoning and expected trade-off.
6. Design for All Users
The interface communicates infrastructure complexity to both business and technical stakeholders, no special knowledge required.
Key screens in the experience:
Dashboard
Cost overview, usage patterns, and optimization opportunities at a glance
Cost Analytics
Detailed cost breakdown, trends, and cost-per-request analysis
Model Comparison
Side-by-side comparison of models with cost, latency, and quality
Optimization Workflow
Guided experience to configure routing and distillation strategies
Routing Configuration
Define model selection rules and automatic optimization strategies
Monitoring
Track impact of optimizations over time with alerts and insights
Experience the interface firsthand:
Explore Distillfast Live →Decisions
Decision 1: From Metrics to Decisions
Instead of presenting infrastructure metrics as the primary experience, organize information around the decisions users need to make. The dashboard asks "what should you do?" not "what are the metrics?"
Decision 2: Cost + Quality + Latency
Cost savings should never be presented without the context of quality and performance. This decision prevented the product from becoming a cost-optimization-at-any-cost tool and ensured users remained aware of the trade-offs they were making.
Decision 3: Progressive Disclosure
Expose high-level insights first and allow technical users to go deeper when needed. This pattern kept the interface approachable for non-ML users while giving power users the control they need.
Decision 4: Make Automation Explainable
If the system recommends a smaller model or routing strategy, users should understand the reasoning and expected trade-off. This builds trust and gives users the confidence to act on recommendations.
Decision 5: Design for Both Technical and Business Users
The interface communicates infrastructure complexity without requiring every user to understand ML architecture. Founders can understand cost, engineers can configure routing, product teams can evaluate trade-offs—all in one product.
Impact
Distillfast was designed to make AI cost optimization a visible product concern rather than an invisible infrastructure problem.
The experience brings together cost visibility, model performance analysis, optimization opportunities, and routing decisions into one coherent workflow. Rather than asking teams to reason about model performance in isolation, the product makes the relationship between usage, cost, and quality immediately visible.
For startups:
Clear understanding of AI infrastructure costs makes product economics more predictable and understandable.
For engineers:
Control over model routing and clear visibility into the impact of optimization decisions.
For product teams:
A shared language for understanding AI economics and making trade-off decisions across technical and business teams.
For users:
A product that works well without knowing the AI infrastructure complexity behind the scenes.
Reflection
"One of the biggest design challenges was translating infrastructure complexity into a product experience that felt understandable without making it simplistic."
AI infrastructure products can easily become dashboards full of numbers. The more interesting UX opportunity is helping users understand what those numbers mean and what they should do next.
Working on Distillfast reinforced several key principles for designing technical products:
Distillfast reinforced the importance of designing not just for visibility, but for decision-making. The best UX isn't about showing users more information—it's about helping them understand what information matters and what to do about it.