Why most AI hackathons lead nowhere, and how to run one that doesn’t.
Many AI hackathons end with demos that vanish. Build the event around real workflows whose owners attend, approved examples, a small evaluation set per team and a path, agreed in advance, from the winning idea to a funded discovery or pilot.
veridive5 min read
The winning team got a trophy, a slot at the all-hands meeting and a round of applause. Its demo, an assistant that answered supplier questions, was genuinely impressive. Nobody asked who owned supplier questions, whether the answers were right on real emails or who would pay for the next step, so nothing happened next.
A hackathon is good at producing energy and demos. It is bad at producing anything that survives the following Monday, unless the event is built for that from the start: real workflows whose owners attend, approved examples, a small evaluation set per team, and a path to a funded next step agreed before anyone writes code.
Why do most AI hackathons lead nowhere?
Because they are designed to produce demos, and demos are the easy part. Four patterns recur, each with a warning sign you can spot in advance:
- The blank canvas. Teams bring their own problems. Warning sign: pitches start with a technology, and no workflow owner is in the room.
- The data free-for-all. Real data travels on laptops. Warning sign: someone asks for “a quick export” of customer records the week before.
- The applause meter. The best presenter wins. Warning sign: the jury sees slides and a live demo, never results on cases the team hasn’t seen.
- The orphan demo. The winner gets a trophy and no next step. Warning sign: nobody can say who funds the follow-up or whose time it needs.
Each has a fix, and every fix is cheap if it is decided before the event.
How do you choose the problems?
Workflow owners bring them, on problem cards prepared in advance. A card names the workflow, the owner, the volume, a set of real examples and what a good outcome looks like, and the owner commits to attending. Screen the cards with the four tests for any first project (frequent, bounded, reachable data, a named owner), or with your use-case prioritization matrix if you have one.
Here is an illustrative example, a filled-in card:
| Field | Illustrative entry |
|---|---|
| Workflow | Supplier emails to the purchasing inbox, sorted into confirmations, delivery changes, price changes, invoices and other, then routed to the right buyer |
| Owner | Purchasing operations lead, present on both days |
| Volume | A few hundred emails a week, more at month-end |
| Examples | Sixty past emails, masked, each with the correct category and buyer; ten held back for judging |
| Definition of good | Right category and buyer on the first pass; delivery changes never missed; unclear emails go to a person |
| Data and limits | Synthetic supplier names, no live contract prices, approved sandbox only |
A card like this takes an owner an afternoon. If nobody will spend that afternoon, the problem isn’t ready for a hackathon.
What data can teams safely use?
Approved or synthetic data only. Prepare masked samples in advance, generate synthetic records where real ones are too sensitive, and run everything in an approved environment with approved tools. No personal accounts, no production write access, no exports to laptops.
Check with your data protection officer before the event, not during it: which data may be used, where it may be processed, and when it will be deleted afterwards. Security should see the environment before the teams do.
How should entries be judged?
On evidence. Every card carries a small held-out set, the ten emails in the example above, that teams never see until judging. Each team runs its entry on those cases and reports results by case type, one failure and why it happened, a rough cost per case, and what the next step would need. The owner’s commitment counts as heavily as the score: would they give time to take it further?
Applause measures the presenter. Held-out cases measure the idea.
Put owners and someone from IT or security on the jury, and use a fixed demo format so the clearest speaker doesn’t win by default.
What happens to the winners?
Whatever was promised before the event. Agree in advance that the top one or two ideas get a committed next step, a budget and protected time for the owner and part of the team. The right next step depends on the evidence: a proof of concept if feasibility is still open, a discovery if the workflow needs a baseline and a plan, a pilot if the results are already strong. The note on proofs of concept, pilots and MVPs helps choose.
The cards that didn’t win go into the use-case backlog with their evidence attached. They are often the next round’s best candidates.
What should the event itself look like?
- Before: cards approved, data prepared and cleared, environment tested, teams formed around cards with at least one person from the owner’s team, and a short primer on evaluation for everyone.
- During: one or two days in working hours; owners present their cards at the start; a midpoint check where every team runs its examples; final demos in a fixed format.
- After: the winners’ next step starts as agreed, what was learned is published, and the data is deleted.
People who will build or use these systems learn more in two days on a real card than in a week of slides, which makes a good hackathon a useful step in a curriculum by role.
Cards before invitations
Draft three problem cards before announcing anything. If you can’t fill in the owner and the examples, the event isn’t ready. The same principles (real workflows, real owners, evidence) run through our AI enablement work, and an Executive Build Day applies them with one leadership team in a single day.
Ask an assistant about this note