Episode 5: The Future of On-Call: Simplicity, AI, and Building Spike with Kaushik Thirthappa

In this episode of The Root Cause, Priyank Upadhyay sits down with Kaushik Thirthappa, founder of Spike.sh, to explore why simplicity is the hardest product problem in tech, the Super Mario approach to SaaS onboarding, status page transparency, and how intent-driven infrastructure will reshape on-call rotations.

Listen on:

If an engineering team has to run an internal training session just to teach their developers how to use an on-call alerting tool, the user experience is broken.

When a production database falls over at 3:00 AM, the human on call is operating in a state of sudden adrenaline and cognitive fog. They do not want a complex command center packed with sixteen nested buttons, obscure status toggles, and multi-tier drop-down menus. They want immediate clarity: What broke, who needs to know, and what is the fastest way to acknowledge it?

In the fifth episode of The Root Cause, host Priyank Upadhyay sat down with Kaushik Thirthappa, founder of Spike.sh. Kaushik is a software developer turned bootstrapped founder whose alerting platform now supports more than 260 engineering and operations teams worldwide.

Their discussion pulls back the curtain on the real friction of building developer tools, the misconception of customer-driven UX, and why the future of site reliability belongs to intent-driven infrastructure.

The "Super Mario" Philosophy of Onboarding

Most enterprise software onboarding follows an identical, passive pattern. You register an account, land on a blank dashboard, and face an empty checklist: connect your cloud provider, configure your escalation policies, invite five team members, and wait.

The core flaw with that workflow in incident response is time-to-value. How does an engineer know your alerting tool actually works if they have to wait two weeks for an actual production outage?

Spike.sh solved this by borrowing a principle from game design: the Super Mario approach.

Classic games do not show players an exhaustive ten-minute tutorial video. World 1-1 is the tutorial. You see a Goomba, you press jump, and you immediately understand the mechanic.

When a team registers on Spike, the platform bypasses the blank slate entirely. It immediately simulates a live test incident and initiates an automated phone call to the user's mobile device, bypassing "Do Not Disturb" settings. Before the user even starts configuring integrations, they have already experienced the exact aha moment of receiving, acknowledging, and resolving an escalation.

Designing that guided experience required intense internal discipline. It meant resisting the temptation to add more onboarding steps and focusing strictly on delivering proof of work within sixty seconds.

Why Customers Should Not Design Your Interface

Founders are constantly told to listen to their users. While user conversations are vital for discovering real underlying pain points, letting customers dictate your interface layout is a recipe for product decay.

Kaushik points out that when you ask operators what they want on an incident table, they will ask for everything: assigned responders, alert counts, timeline summaries, runbook links, and direct remediation buttons. If you follow every request, your dashboard ends up resembling an air traffic control console.

"Making something simple is so much harder than making it complicated," Kaushik explains.

Customers give feedback based on existing touchpoints and past habits. They can tell you where friction exists, such as needing an out-of-office override because a teammate is away for a child's birthday. But the burden of translating that need into an intuitive, minimal layout rests solely on the product team. Simplicity requires aggressive pruning.

The New Zealand Warning Sign: Radical Transparency

A pivotal moment in Kaushik's philosophy on reliability came during a family vacation to New Zealand. Shortly before arriving in Queenstown, severe brushfires broke out across the region, cutting off major travel routes toward Milford Sound.

While visiting a scenic river roughly sixty kilometers outside the town, Kaushik walked down to the water. Instead of seeing a hastily taped piece of paper or an improvised barricade, he was greeted by a polished, newly installed wooden board explaining that the water was temporarily contaminated due to the recent fire runoff.

The response demonstrated incredible institutional preparedness. The local authorities had pre-manufactured high-grade warning signage and deployed it across dozens of remote tourist destinations within twenty-four hours.

That level of transparency left a permanent mark on how Spike approached incident communication.

When an infrastructure failure occurs, many engineering teams hesitate to update their public status page. The internal debate is always driven by fear: Will customers lose faith in our software? Will prospects walk away?

In reality, transparency has the opposite effect. Engineering teams do not lose trust because an outage happens; they lose trust when a platform behaves evasively. Platforms like GitHub, Linear, and leading frontier AI labs post status updates rapidly and objectively. Adopting an open, blameless status page policy turns an operational failure into a demonstration of accountability.

The 10-Year Horizon: From Alert Firefighting to Intent-Driven Systems

Toward the end of the conversation, Priyank and Kaushik tackled the long-term evolution of the Site Reliability Engineer.

Over the past decade, operations teams have been saddled with an unsustainable burden: context-switching across monitoring tabs, triaging thousands of noisy alerts, and resolving routine configuration drift during off-hours.

Priyank offered a contrarian perspective on where this lands in the next five to ten years: infrastructure will eventually become self-healing and intent-driven.

Consider how developers interact with compilers today. When writing high-level code, programmers do not spend mental bandwidth worrying about how CPU registers are allocated or how machine instructions are organized. The compiler abstracts the execution.

Infrastructure is heading toward a similar abstraction layer. Rather than manually wiring alerts and writing custom bash scripts to restart crashed pods, operators will simply declare system intent: "Keep this service accessible, keep latency under 100 milliseconds, and isolate anomalies."

Autonomous reliability engines will handle state reconciliation in the background. The role of the engineer will shift left, migrating away from 3:00 AM tactical firefighting and elevating into high-level business architecture and intent verification.

See how it works.

Book a 30-minute demo. No slides, just your stack.

Download Whitepaper