05 Field notes
Good questions. Practical thinking.
Ideas for people putting AI to work — from the projects we run and the mistakes we’ve learned from.
01 Featured
Start with these.
-
Choose the workflow before the model.
Most stalled AI projects didn’t fail on technology. They failed because nobody picked a specific piece of work to change. Here’s how we choose — and what we measure before writing any code.
-
What makes an AI pilot ready for production?
A pilot proves something is possible. Production proves it’s reliable, affordable and owned. Our checklist covers evaluation thresholds, approval points, cost ceilings, monitoring and the runbook your team will actually use.
-
An AI glossary for business teams, in plain words.
The terms you hear in every vendor meeting, each defined in two plain sentences with the one question it should make you ask. Grouped by the decision each term affects, not by the alphabet.
-
RAG, explained for business teams: how AI answers from your documents.
Retrieval-augmented generation is less about the model than about the library: which documents are approved, who may see them, how they are found and how every answer points back to its source. Most quality problems start there.
-
KVKK, GDPR and AI: questions to ask your DPO before you build.
In Türkiye and Europe, data protection is often the first objection to an AI project. It becomes answerable once the data flow is mapped. These are the questions to settle with your DPO or counsel before anyone builds.
-
Build, buy or wait: how to decide for each AI workflow.
Decide per workflow, not per company. Buy what every company needs, build where the workflow is how you compete or where you must control data, prompts and evaluation, and wait when there is no owner or no reachable data.
02 All notes
Every note, by topic.
100 notes
-
Choose the workflow before the model.
Most stalled AI projects didn’t fail on technology. They failed because nobody picked a specific piece of work to change. Here’s how we choose — and what we measure before writing any code.
-
What makes an AI pilot ready for production?
A pilot proves something is possible. Production proves it’s reliable, affordable and owned. Our checklist covers evaluation thresholds, approval points, cost ceilings, monitoring and the runbook your team will actually use.
-
Evaluation sets are the new requirements document.
When the output is language, “done” is hard to define. A few hundred real examples with expert-approved answers make quality measurable — and turn model choice into an experiment instead of an opinion.
-
A useful system needs a new everyday practice.
Adoption doesn’t come from a launch email. It comes from new routines, champions who help colleagues, and managers who ask for the new output in their weekly meetings.
-
Where agents should — and shouldn’t — act on their own.
Agents are powerful when actions are reversible, inexpensive and well logged. For everything else, they should prepare and a person should decide.
-
The cost of a token vs. the cost of a mistake.
Model prices keep falling; mistakes don’t get cheaper. We compare cost per task with the cost of an error — and often end up with a small model for most cases and a person for the rest.
-
Answers with receipts: why citations matter at work.
Before veridive helped companies, it answered questions from videos with a citation to the exact second. The lesson carried over: people trust answers they can check.
-
Questions leadership teams ask before their first AI project.
Where do we start? What will it cost to run? What happens to our people? Who’s accountable when it’s wrong? Our answers to the questions we hear most.
-
An AI glossary for business teams, in plain words.
The terms you hear in every vendor meeting, each defined in two plain sentences with the one question it should make you ask. Grouped by the decision each term affects, not by the alphabet.
-
RAG, explained for business teams: how AI answers from your documents.
Retrieval-augmented generation is less about the model than about the library: which documents are approved, who may see them, how they are found and how every answer points back to its source. Most quality problems start there.
-
KVKK, GDPR and AI: questions to ask your DPO before you build.
In Türkiye and Europe, data protection is often the first objection to an AI project. It becomes answerable once the data flow is mapped. These are the questions to settle with your DPO or counsel before anyone builds.
-
Chatbot or agent assist: where should customer service AI start?
Start with AI that helps agents: suggested replies with sources, summaries and routing. Agents catch the mistakes while you learn which requests are safe to automate, then open self-service for those few, with a person always one step away.
-
Build, buy or wait: how to decide for each AI workflow.
Decide per workflow, not per company. Buy what every company needs, build where the workflow is how you compete or where you must control data, prompts and evaluation, and wait when there is no owner or no reachable data.
-
What makes Turkish hard for AI, and how to design around it.
Turkish builds meaning with suffixes, has a dotted and dotless i that breaks naive software, and often shares a sentence with English. None of that stops a good system; it means testing on real Turkish text and designing search, casing and tone on purpose.
-
Invoice automation in Türkiye: where e-Fatura ends and AI begins.
Structured e-invoices already arrive as data. The hard part of invoice automation in Türkiye is everything else: foreign supplier PDFs, scans, delivery notes, credit notes and matching to orders. That is where AI helps, with an accountant approving every entry.
-
How to write an AI business case that finance will approve.
A credible AI business case prices the whole workflow, including running costs, review time and errors, against a measured baseline. It writes down the assumptions finance can check, and names the point at which you stop.
-
Proof of concept, pilot or MVP: which AI step are you taking?
A proof of concept asks whether a model can do the task, a pilot asks whether it works on live work, and production promises it will keep working. Confusing the three is how impressive demos die.
-
Why AI makes things up, and how to design so it gets caught.
You can’t remove hallucinations entirely, but you can make them rarer and visible: ground answers in approved sources, allow “I don’t know”, check names and numbers against systems of record, and put a person where a wrong answer is expensive.
-
Do you need an AI agent, or a workflow with a few model calls?
Most business processes are better served by a fixed workflow with model calls at specific steps. Agents earn their place where the path can’t be known in advance, and they cost more to test, run and explain.
-
What to monitor after an AI system goes live, and why quality slips.
Quality slips without anyone touching the system: sources change, the case mix shifts, providers update models. Watch quality, cost, speed and use together, and treat a rising override rate as the earliest warning you will get.
-
AI literacy: what every employee should know before using AI at work.
AI literacy isn’t knowing how models work inside; it is knowing what they are good and bad at, which data must never go in, how to check an answer and when a person decides. All of it fits on one page.
-
Shadow AI: what to do when employees already use their own tools.
Banning AI tools pushes their use into personal accounts you can’t see. Find out what people use and why, offer an approved tool that is genuinely good, set clear data rules and turn the best unofficial uses into projects.
-
When not to use AI: signs a rule, a form or a report will do better.
Many AI requests are better solved by a rule, a lookup, a redesigned form or a report. Where the answer must be exact, the volume is tiny or nobody owns the outcome, a model adds cost and risk without adding value.
-
RAG or fine-tuning: which one does your use case need?
Retrieval gives a model knowledge it can cite and you can update daily; fine-tuning changes how it behaves. Most business knowledge problems need retrieval first, and fine-tuning only when evaluation shows a gap prompts can’t close.
-
How to prioritize AI use cases when the list has thirty ideas.
Score only what you can evidence: volume, time per case, cost of errors, reachable data and a named owner. Keep value and feasibility on separate axes, because one weighted score hides the trade-off leadership actually has to make.
-
Prompt injection: which defenses actually hold.
You can’t filter your way out of prompt injection. Limit what an injected instruction could achieve instead: separate instructions from data, scope permissions, validate outputs and put a person in front of consequential actions.
-
How to measure AI in customer service beyond handle time.
Average handle time can fall while customers call back. Measure resolution, repeat contacts, the accuracy of cited answers, how much agents edit drafts, escalations and satisfaction by request type, against the baseline recorded before launch.
-
Hours saved are not money saved: measuring the ROI of an AI system.
Time saved becomes value only when someone decides what the time is for. Measure against the baseline, attribute results carefully, and count the quality, speed and risk gains the business already puts a price on.
-
Do you need an AI assistant, or just better search?
When people need to find a document, better search with good metadata may be enough, and it costs less. When they need an answer assembled from several passages, with a citation, an assistant earns its keep. Many teams need search first.
-
What to teach each role about AI, from the board to the front line.
One course for everyone teaches that the tool exists. A curriculum by role teaches when it helps, where its limits are and what each person is now expected to do differently.
-
Running your own model or using an API: how to decide.
Self-hosting an open-weight model makes sense when data must not leave your infrastructure, volumes are high and steady, or you need control over change. The price is operations. Decide per workflow with the evaluation set, not per company by principle.
-
Shadow mode: how to test AI on live work without risking it.
Run the system on real cases while people work as usual, compare its proposals with what they decided, and let the agreement rate by case type decide what moves forward. It is the cheapest honest test there is.
-
AI for finance teams: from invoices to month-end, where to start.
Finance work is full of reading, matching and explaining, which suits AI, and full of numbers that must be exact, which doesn’t. Let systems compute and reconcile, let AI read and draft, and keep approvals with the people who sign.
-
Designing AI for teams that work in Turkish and English.
In many organizations the question arrives in Turkish, the policy is in English and the customer writes in both. Decide which language answers go out in, search across languages, cite the source in its original language and evaluate in both.
-
When AI gets it wrong: an incident plan for the first hours.
AI incidents are rarely outages. They are wrong answers that look right. Set severity by impact, make the stop switch and fallback routine, tell affected people quickly, and turn every incident into new evaluation cases.
-
Where does your data go when you use a language model?
Data travels further than people think: prompts, retrieved passages, outputs, logs, caches, search indexes and test sets. Map each hop, check the contract and settings for each, and design so the most sensitive data never makes the trip.
-
Is your company ready for AI? Ask one workflow at a time.
Company-wide maturity scores rarely predict success. Readiness belongs to a workflow: reachable data, a named owner, a baseline you can measure, an error you can tolerate and people with time to take part.
-
What are AI guardrails? A map of where each one lives.
“Guardrails” covers very different controls: checks before the model sees a request, limits on what the system can do, validation of what it produces, and people at approval points. The most dependable are ordinary software, not instructions in a prompt.
-
WhatsApp customer service with AI: what to automate and what not to.
Customers treat WhatsApp like a conversation with a person, so automation has to be quick, honest about being automated and one message away from a human. Automate lookups such as order status first; keep complaints, payment problems and anything sensitive with people.
-
What drives the cost of an AI project, and how to keep it predictable.
The model is rarely the expensive part. Cost follows the systems you connect, the state of the data, the review design, the languages, the security reviews and how much routines change. A fixed scope per phase keeps it predictable.
-
Before the AI assistant: a clean-up checklist for your documents.
An assistant can’t be more current or consistent than its library. Before building, name an owner for each document set, retire duplicates and old versions, fix what scans and tables hide, and mark what is draft, final or expired.
-
What to tell your teams about AI, and how to say it.
People fill silence with the worst case. Say early and specifically which tasks change, which decisions stay with people, what the time saved is for and what training comes, and give managers words for the questions they will get.
-
Why keyword search still matters in AI retrieval.
Vector search finds meaning; keyword search finds exact things such as part numbers, invoice numbers, legal terms and names. Business questions need both, plus reranking, and Turkish word forms make the keyword side harder and more important.
-
How to choose an AI partner: questions that show how they really work.
The answers that matter are about evidence and ownership: whether a firm can show an evaluation set from past work, who will actually build, what it refuses to automate, how it estimates running costs and what you own when it leaves.
-
Four ways to connect AI to your ERP, and the trade-offs of each.
There are four practical ways in: the ERP’s own APIs, read-only database views, scheduled exports, and middleware or screen automation as a last resort. Choose per workflow by freshness, safety and upkeep, and start read-only.
-
Reading Turkish documents with AI: OCR, tables, stamps and signatures.
Most document errors start before the model: a scanner that reads ş as s, a table flattened into a paragraph, a stamp over an amount. Route each document type deliberately, measure errors on Turkish characters and key fields, and send unreadable pages to a person.
-
How to switch language models without breaking what works.
Models get retired, prices change and better options appear, so plan to switch from day one. With an evaluation set, a model-agnostic layer and side-by-side runs, a migration becomes a test result, not a leap of faith.
-
An AI acceptable use policy people will actually read.
A policy that fits on two pages and answers real questions, which tools, which data, who checks and what to disclose, gets followed. A long list of prohibitions gets ignored and pushes use into personal accounts.
-
Productivity assistants or custom AI systems: what is each one for?
General assistants make individuals faster at drafting, summarizing and searching. Custom systems change a workflow, with access to its data, approval points and measurement against a baseline. You will probably need both, bought and built for different jobs.
-
How to design a review screen people can decide from in one look.
The review screen decides whether AI saves time. Show the recommendation, the evidence behind it, what is uncertain and what will happen on approval, so a person can decide without redoing the work.
-
Handing a customer from bot to person without making them repeat it.
The handover decides how customers remember the automation. Escalate on clear triggers, pass a summary with the data already gathered, and tell the customer honestly what happens next.
-
What is a token? The unit behind AI costs and limits, explained.
Tokens are how language models read, write and bill. They explain context limits, why long documents cost more, why output is priced differently from input, and why the same text can cost differently across languages and models.
-
Why document assistants give wrong answers, and where to look first.
When a document assistant is wrong, the model is usually the last suspect. Most errors come from the library and the search: outdated or duplicate documents, broken tables, passages cut mid-thought, missing permissions or a question that needed clarifying.
-
Better prompts at work: context, examples and a way to check.
A good prompt is a good brief: context, the goal, an example of what good looks like, the format you need and how the answer will be checked. There are no magic words, and the best prompts become shared templates.
-
How to make an AI assistant respect who may see what.
An assistant must never become a way around your permissions. Enforce access at retrieval time from the source system’s rules, keep those rules in sync, and test with real personas, including people who just lost access.
-
How to write an AI RFP that gets answers you can compare.
Describe the workflow, real examples and what good looks like instead of a feature list. Ask every bidder how they would evaluate, who would build, what you would own and what it costs to run, then decide with a short paid test.
-
Management reports with AI: numbers from systems, words from the model.
Never let a language model calculate the numbers in a management report. Compute figures in the ERP or BI layer, let the model draft commentary from them and the notes people already write, and cite every statement to its figure.
-
Turkish speech-to-text for calls and meetings: what affects accuracy.
Transcription quality decides everything downstream in call and meeting intelligence. Errors cluster in phone audio, accents, crosstalk, numbers, names and English terms inside Turkish speech; measure on your own recordings and design review around the words that matter.
-
Every override is a lesson: building a feedback loop that improves AI.
Thumbs-up buttons collect noise. Structured feedback from the review step (what was changed and why), triaged weekly by the owner, turns daily corrections into better sources, prompts and evaluation cases.
-
What an AI audit trail should record, and who should read it.
When a decision is questioned months later, you need to rebuild what the system saw, which sources it used, which model and prompt answered, and who approved it. Log that by design, and treat the logs as sensitive data.
-
What belongs in an AI roadmap, and what should stay out of it.
A useful AI roadmap is a sequence of workflows, each with an owner, a baseline and a go/no-go gate, arranged as now, next and later rather than by calendar quarter. Technology appears only where a workflow needs it.
-
Getting reliable structured data out of a language model.
Treat model output like any untrusted input: define a strict schema, constrain generation, validate every field against business rules, and send what fails to a person instead of retrying forever.
-
How to design an intent taxonomy that AI routing can live with.
Routing fails more often because of the categories than because of the model. Fewer, mutually exclusive intents, each with examples and an owner, make AI routing measurable.
-
How to forecast LLM running costs before the first invoice arrives.
Forecast from the task, not from a price list: calls per task, tokens per call, retries, volume peaks and review minutes, all measured in the pilot. Then set a ceiling with alerts before launch.
-
Clause extraction for legal teams: what AI can and can’t check.
AI is good at finding, extracting and comparing clauses against your playbook across many contracts; it is not the lawyer. Start with a clause list and a playbook your legal team wrote, cite every extraction, and keep judgment on risk with people.
-
What your software engineers need to learn to build AI systems.
Good software engineers already have most of what it takes. The new parts are an evaluation mindset, output that varies between runs, retrieval and data handling, and judgment about where a person decides, learned fastest by shipping one workflow with someone who has done it.
-
An evaluation set template: columns, tags and a scoring rubric.
The working sheet behind an evaluation set: the columns, tag lists, rubric wording and splits to start from, described so a team can build it in an afternoon, with filled-in rows from a returns workflow and a policy assistant.
-
What your team should receive when an AI project ends.
A handover is complete when your team can change a prompt, re-run the evaluation set, roll back a model and explain a decision to an auditor without calling the supplier.
-
What language models add to demand planning, and what they don’t.
Forecasts are a job for statistical and machine learning models on sales history. Language models help around them: reading supplier and customer messages for signals, explaining forecast changes in plain words and answering planners’ questions with sources.
-
AI translation at work: when it’s good enough and when it isn’t.
AI translation is good enough for understanding and internal drafts, and risky for contracts, regulated text and anything published under your name without review. A glossary, a reviewer for high-stakes text and a clear rule on what may be sent where make the difference.
-
Why reviewers stop checking AI output, and how to keep review real.
When a system is right most of the time, people stop looking. Evidence-first screens, known-answer cases, rotation and metrics that notice when approval becomes a reflex keep human oversight meaningful.
-
Questions to ask an AI vendor before you sign.
Standard security questionnaires miss what matters for AI products: which models sit underneath and who can change them, whether your data trains anything, how quality is measured on your cases, and what you can take with you when you leave.
-
Generative AI or machine learning: which does your problem need?
Forecasting, scoring and anomaly detection on structured data are usually classic machine learning. Reading, writing and reasoning over documents and conversations is where language models help. Many good systems use both, each for the part it does well.
-
How to test an LLM application, layer by layer.
Test the deterministic parts like ordinary software, the language parts with evaluation sets, and the joins between them. A single end-to-end accuracy number hides where things actually break.
-
Product descriptions with AI that don’t all sound the same.
Generated descriptions go wrong in two ways: invented features and identical copy across thousands of products. Write only from verified attributes, vary structure by category, list the claims the model may never make, and let merchandisers approve what is published.
-
How to cut LLM costs without quietly cutting quality.
Most savings come from sending less, and sending it to the right place. Every change is checked against the evaluation set, because a cheaper system that needs more review is not cheaper.
-
How to build a call quality scorecard AI can apply consistently.
AI can review every call only against criteria it can apply the same way twice. Rewrite vague criteria such as “was empathetic” into observable behaviors, calibrate on past calls with the quality team, and cite every finding to the second.
-
Why most AI hackathons lead nowhere, and how to run one that doesn’t.
Many AI hackathons end with demos that vanish. Build the event around real workflows whose owners attend, approved examples, a small evaluation set per team and a path, agreed in advance, from the winning idea to a funded discovery or pilot.
-
When can you trust an AI to grade another AI’s answers?
Automated grading makes evaluation cheap enough to run on every change, but the judge is one more system to validate. Use narrow yes/no checks, calibrate against people, watch for known biases and keep people on the high-risk cases.
-
From one team to many: how to roll out an AI system that worked once.
The second team is never a copy of the first: its cases, language, systems and habits differ. Re-baseline, re-run the evaluation set on its cases, find local champions and widen scope one step at a time.
-
AI for procurement: spend, supplier documents and quote comparison.
Procurement is heavy on documents and on judgment. AI can classify spend, read supplier documents and quotes, and compare offers line by line with sources; the choice of supplier and every commitment stay with buyers.
-
Arabic and English in one AI system: what to plan for in the Gulf.
In the Gulf, work moves between Arabic and English, and Arabic itself spans formal written Arabic, dialects and Latin-script chat. Plan for script and direction, name spellings across scripts and native reviewers, and test quality per language instead of assuming it.
-
LLMOps and MLOps: what changes when the model is a language model?
Much of MLOps carries over, but with language models you rarely train: you manage prompts, retrieval sources, external model versions and human review. The center of gravity moves from training pipelines to evaluation and change control.
-
The EU AI Act: questions to ask if you build or buy AI for Europe.
The EU AI Act sorts AI uses by risk and gives duties to providers and deployers. Settle the basics with counsel first: which role you play, which category each use falls into, and what transparency and oversight you already have.
-
What an AI center of excellence should own, and what it shouldn’t.
A small central team should own shared foundations, standards, evaluation practice and vendor contracts. Workflows and their results belong to business owners. A center of excellence that approves every idea turns into a queue.
-
Prompts are code: how to version, review and test them.
A prompt change is a behavior change. Keep prompts in version control, review them like code, tie each version to evaluation results, and never edit them live in a dashboard.
-
Mapping your catalog to every marketplace’s categories with AI.
Each marketplace has its own category tree and required attributes, and a wrong category hides a product or gets it rejected. AI can propose categories and fill attributes from your data, while people approve each new mapping once.
-
Meeting minutes with AI: decisions, owners and a link to the moment.
A useful meeting record is not a prose summary. It lists decisions, owners, deadlines and open questions, each linked to the moment it was said, drafted by AI and approved by a person, with consent and retention agreed up front.
-
Asking your database in plain language: where text-to-SQL works.
Text-to-SQL works on a small, documented set of curated views with business definitions and read-only access; pointed at raw ERP tables, it produces confident, wrong numbers. Show the query and the data behind every answer.
-
Who you need on an AI project, and why the domain owner matters most.
A small team with a domain owner, engineers who can evaluate as well as build, and a designer for the review step beats a large team of specialists. The scarcest resource is the domain owner’s time, so plan it first.
-
Logistics documents and AI: what to automate first.
Logistics runs on documents from many parties, in many formats and languages. Start where a document triggers a decision, such as a proof of delivery that releases an invoice, extract with evidence and route exceptions to the right desk.
-
What a support agreement for an AI system should promise.
You can’t promise accuracy the way you promise uptime. A useful agreement commits to quality thresholds on the evaluation set, incident response times, rules for model and prompt changes, cost reporting and a regular improvement review.
-
Red-teaming an AI system: a practical test plan before launch.
Red-teaming means structured attempts to make the system misbehave: leak data, skip an approval, give off-policy answers or run up costs. Do it with people who know the business, log every finding, and turn each one into a test.
-
What should a board ask about AI? Questions for directors.
Boards don’t need model details. They need a register of AI systems in production with named owners, the decisions each one can influence, how quality and cost are measured, and what happened in the last incident.
-
How to design the tools an AI agent is allowed to use.
An agent is only as safe and reliable as its tools. Make each tool narrow, clearly described, scoped to minimum permissions, safe to call twice and informative when it fails.
-
Making AI replies sound like your brand in Turkish and English.
Models follow examples better than adjectives. A tone guide for AI replies needs rules a model can apply (form of address, sentence length, what never to say, how to apologize), paired examples in each language and checks in the evaluation set.
-
Masking personal data before it reaches a language model.
Masking reduces what a model provider sees, but it is not anonymization and it fails quietly. Detect locally, use reversible placeholders where the answer needs them, test detection on your own data formats, and don’t treat masking as the whole privacy plan.
-
AI for tender and RFP responses: a playbook for bid teams.
Bid teams answer the same questions again and again, under deadline. AI helps most by finding the best past answer with its source, drafting for the new question and flagging what has changed, while subject experts approve every commitment.
-
Cleaning supplier and material master data with AI, safely.
Duplicate suppliers and inconsistent material descriptions quietly break matching, reporting and every AI workflow built on top. AI can propose merges and standard descriptions with evidence; data owners approve each change, and nothing merges automatically.
-
AI in a group of companies: what to share and what to leave local.
In a group, share the foundations: model access, security patterns, evaluation practice, contracts and training. Let each company choose and own its workflows. A group-wide platform mandate before any workflow works usually slows everyone down.
03 By email
Field notes in your inbox.
Occasional field notes by email. One practical idea at a time.