DEV Community

Debashish Ghosal profile picture

Debashish Ghosal

TLDR - Engineer -> Manager -> Tech Leadership, passionate about AI, Developer experience, planet scale infra, transformation - AI-native leadership. Rest here - https://www.linkedin.com/in/deghosal/

Location Seattle, WA Joined Joined on  Personal website https://github.com/deghosal-2026/ github website

Pronouns

he/him

Work

Engineering Leadership

I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.

I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.

6
Comments
7 min read

Want to connect with Debashish Ghosal?

Create an account to connect with Debashish Ghosal. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
The Same Model Debating Itself Was More Self-Critical Than Two Different Models

The Same Model Debating Itself Was More Self-Critical Than Two Different Models

13
Comments
13 min read
Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong

Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong

16
Comments 2
10 min read
The Best Model Pair in My Field Test Was Also the Least Trustworthy

The Best Model Pair in My Field Test Was Also the Least Trustworthy

21
Comments 7
12 min read
I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.

I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.

14
Comments
31 min read
Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.

Reveals why blind reviews fail

Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.

13
Comments 8
14 min read
My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.

My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.

21
Comments 5
5 min read
My Agent Refused 96 Times. That Was the Right Output.

My Agent Refused 96 Times. That Was the Right Output.

22
Comments 5
10 min read
Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

Tested on 70 real PRs for just $0.53 total cost

Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

16
Comments 12
11 min read
A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.

A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.

18
Comments 8
7 min read
I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.

Structural gates beat prompt safety

I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.

40
Comments 9
12 min read
I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

17
Comments 5
13 min read
The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.

Exposes the three recurring failure patterns

The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.

21
Comments 28
5 min read
I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'

Code-enforced allowlists fix prompt drift

I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'

16
Comments 26
4 min read
I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.

Explores why self-review fails agents

I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.

33
Comments 32
8 min read
AI Governance Is Becoming a Transformation Problem

AI Governance Is Becoming a Transformation Problem

22
Comments 4
8 min read
Your Company Has AI Tribes. Send an Engineer as Emissary

Adapting Palantir's model internally

Your Company Has AI Tribes. Send an Engineer as Emissary

7
Comments 4
12 min read
I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.

I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.

8
Comments 2
13 min read
AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

13
Comments 3
13 min read
I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.

Adapters need conformance suites too

I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.

18
Comments 21
15 min read
I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

Real production data broke 35 detectors instantly

I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

22
Comments 19
8 min read
Why Agent Evaluation Is Harder Than Model Evaluation

Builder scars reveal untrustworthy agent paths

Why Agent Evaluation Is Harder Than Model Evaluation

19
Comments 25
8 min read
Faster PRs, Weaker Instincts: The Judgment Problem in AI-Assisted Engineering

Metrics mask silent logic gaps

Faster PRs, Weaker Instincts: The Judgment Problem in AI-Assisted Engineering

8
Comments 3
4 min read
AI-Assisted Engineering: Faster to Build Isn't Cheaper to Own

QA and exception handling get denser

AI-Assisted Engineering: Faster to Build Isn't Cheaper to Own

13
Comments 8
6 min read
I Planned 10 LLM Evaluation Experiments And Only Ran 1. It Was Enough.

I Planned 10 LLM Evaluation Experiments And Only Ran 1. It Was Enough.

5
Comments 2
13 min read
I Connected 3 MCP Servers to One Agent. It Got Scary Fast.

Exposes glaring security gaps in multi-MCP setups

I Connected 3 MCP Servers to One Agent. It Got Scary Fast.

15
Comments 33
5 min read
Platform Strategy That Ships: The 90-Day System I Trust More Than AI Roadmaps

Platform Strategy That Ships: The 90-Day System I Trust More Than AI Roadmaps

5
Comments
7 min read
4 Open-Source AI Tools, 1 MCP Server — What I Built and What I Learned

Overcoming internal dev tool adoption hurdles

4 Open-Source AI Tools, 1 MCP Server — What I Built and What I Learned

10
Comments 17
15 min read
The Best AI Productivity Win I Got This Month Wasn't From Coding

The Best AI Productivity Win I Got This Month Wasn't From Coding

Comments
4 min read
Your Flaky Tests Aren't Flaky

Your Flaky Tests Aren't Flaky

Comments
6 min read
i built a tool that tracks what AI tasks actually cost. the real number surprised me.

i built a tool that tracks what AI tasks actually cost. the real number surprised me.

Comments 3
6 min read
i've been building platforms first for 25 years. i think it's wrong now.

i've been building platforms first for 25 years. i think it's wrong now.

1
Comments
6 min read
i tested an ai incident commander against 15 real outages — 88% pass rate

i tested an ai incident commander against 15 real outages — 88% pass rate

1
Comments 2
4 min read
i spent 3 days on a rag pipeline and the embedding model didn't matter at all

i spent 3 days on a rag pipeline and the embedding model didn't matter at all

Comments
6 min read
my ai agent ran for 6 hours on a 2-minute task and cost me $200

my ai agent ran for 6 hours on a 2-minute task and cost me $200

1
Comments 1
6 min read
i've been doing sdlc for 25 years. ai didn't kill it. it made it the coordination layer.

i've been doing sdlc for 25 years. ai didn't kill it. it made it the coordination layer.

1
Comments 1
7 min read
i was using a frontier model to update jira tickets. here's how i fixed it.

i was using a frontier model to update jira tickets. here's how i fixed it.

2
Comments
6 min read
i built two langgraph projects and learned that agents are junior engineers with amnesia

i built two langgraph projects and learned that agents are junior engineers with amnesia

2
Comments 8
7 min read
loading...