# Data & models.

> The layers under a useful AI system: RAG and search, permissions, model and hosting choices, evaluation methods and masking personal data before it travels.

Good answers need the right data, the right permissions and the right model, in that order. These notes explain retrieval, model and hosting choices and evaluation methods in terms IT and data teams can act on.

- [RAG, explained for business teams: how AI answers from your documents](https://veridive.com/insights/rag-explained-for-business-teams/): Retrieval-augmented generation is less about the model than about the library: which documents are approved, who may see them, how they are found and how every answer points back to its source. Most quality problems start there.
- [RAG or fine-tuning: which one does your use case need?](https://veridive.com/insights/rag-vs-fine-tuning/): Retrieval gives a model knowledge it can cite and you can update daily; fine-tuning changes how it behaves. Most business knowledge problems need retrieval first, and fine-tuning only when evaluation shows a gap prompts can’t close.
- [Running your own model or using an API: how to decide](https://veridive.com/insights/self-hosted-llm-vs-api/): Self-hosting an open-weight model makes sense when data must not leave your infrastructure, volumes are high and steady, or you need control over change. The price is operations. Decide per workflow with the evaluation set, not per company by principle.
- [Why keyword search still matters in AI retrieval](https://veridive.com/insights/hybrid-search-rag/): Vector search finds meaning; keyword search finds exact things such as part numbers, invoice numbers, legal terms and names. Business questions need both, plus reranking, and Turkish word forms make the keyword side harder and more important.
- [How to make an AI assistant respect who may see what](https://veridive.com/insights/rag-access-control/): An assistant must never become a way around your permissions. Enforce access at retrieval time from the source system’s rules, keep those rules in sync, and test with real personas, including people who just lost access.
- [An evaluation set template: columns, tags and a scoring rubric](https://veridive.com/insights/llm-evaluation-template/): The working sheet behind an evaluation set: the columns, tag lists, rubric wording and splits to start from, described so a team can build it in an afternoon, with filled-in rows from a returns workflow and a policy assistant.
- [When can you trust an AI to grade another AI’s answers?](https://veridive.com/insights/llm-as-a-judge/): Automated grading makes evaluation cheap enough to run on every change, but the judge is one more system to validate. Use narrow yes/no checks, calibrate against people, watch for known biases and keep people on the high-risk cases.
- [Asking your database in plain language: where text-to-SQL works](https://veridive.com/insights/text-to-sql-for-business-data/): Text-to-SQL works on a small, documented set of curated views with business definitions and read-only access; pointed at raw ERP tables, it produces confident, wrong numbers. Show the query and the data behind every answer.
- [Masking personal data before it reaches a language model](https://veridive.com/insights/pii-redaction-for-llms/): Masking reduces what a model provider sees, but it is not anonymization and it fails quietly. Detect locally, use reversible placeholders where the answer needs them, test detection on your own data formats, and don’t treat masking as the whole privacy plan.
