CORNELL TECH × COLUMBIA · Empire Hacks 2026 · March 2026 1st Place — Best Hackathon Submission

Glassbox

Clinical AI reasoning made visible. The AI forks at uncertainty. The doctor sees how evidence backs every claim — and how different sources converge into each rationale.

Franklin Dickinson · Cole Donat · Alex Seo
Scroll

The reasoning tree, live

The tree grows node by node, auto-pauses at decision points, and lets you click into any branch.

Keep reading

The problem

Clinical AI is a black box

Current AI diagnostic tools return a single answer. Doctors can't see which evidence supports it, how different sources converge, or where the reasoning actually holds up.

01 / VISIBILITY

One answer, no trace

Standard chain-of-thought gives you a diagnosis. It doesn't show you the differential it discarded, the evidence it weighted, or where it was genuinely uncertain.

02 / AGENCY

No alternative paths

A doctor who suspects a different hypothesis has nowhere to look. The model committed to one branch. Everything else is gone.

03 / EVIDENCE

No view into the evidence

Doctors see the conclusion, not what's underneath it. They can't trace a claim back to its sources, or see where the model leaned too hard on one piece.

Transparency built into the reasoning

Glassbox doesn't explain AI decisions after the fact. The tree structure is the reasoning — live, navigable, and interactive.

01

Entropy branching

When the differential is genuinely ambiguous, the backend measures entropy over the model's output distribution and forks into parallel reasoning beams rather than committing to one path.

entropy_branching.py
02

Interactive reasoning tree

The tree grows live and auto-pauses at decision points. Click a branch to isolate it — everything else fades to 20% opacity. Arrow keys and a scrubber let you walk each step with an audience.

react + d3-hierarchy
03

Clinical synthesis

The right panel translates the tree into a structured clinical narrative — each claim traced back to the evidence that supports it, with caveats surfaced where sources diverge.

doctor-in-the-loop

The story

Where it started

The insight

Anthropic's interpretability research — Tracing the Thoughts of a Large Language Model — showed that LLMs reason through traceable computational circuits, run parallel pathways simultaneously, and sometimes fabricate plausible-sounding steps when genuinely uncertain.

If those internal paths can be traced, the question is obvious: why isn't that what a doctor sees?

The brief

Empire Hacks 2026's challenge track was The Auditor: Regulated Agents for Trust — build for a domain where mistakes are expensive. The question judges would use: "Could I audit this agent's work and trust it?"

Traceability weighted 35% of the score. The reasoning chain had to be verifiable at every step — not summarized after the fact, but auditable live. Glassbox's answer: the tree is the audit trail.

Built with
React 18 All UI components — the tree viewport, synthesis panel, growth playback controls, and focus system. Functional components with hooks throughout; no class components.
TypeScript Strict typing across the full frontend: tree node interfaces, the focus state machine, growth playback state, and the transformer layer that converts raw backend data to positioned layout.
Tailwind CSS Primary styling system — all layout, spacing, and typography via utility classes co-located with markup. CSS custom properties handle the clinical color tokens (blue = reasoning, amber = decision, green = tool call).
Framer Motion Node entrance animations and the incremental tree growth effect. Each node fades and scales in as it's revealed during playback. Branch connections animate in sync with their target node.
d3-hierarchy Layout math only — computes node positions and bezier path coordinates for the tree. React renders all SVG directly; d3 never touches the DOM. Clean separation between layout and rendering.
Vite Build tool and dev server. Sub-second HMR made rapid iteration possible during the 48-hour hackathon window without ever waiting on a build.
Python multibeam_search.py runs parallel chain-of-thought beams over the clinical scenario. entropy_branching.py measures output distribution entropy to detect and trigger decision points.
DeepSeek LLM powering the parallel reasoning beams. Each branch is an independent chain-of-thought that explores a distinct diagnostic hypothesis — cardiac vs. GI, for example — without knowledge of the other branches.
Vercel Frontend deployment. Zero-config, instant preview URLs for sharing the live demo with judges and teammates throughout the hackathon.