The pipeline
Clean, chunk, embed, store (often pgvector). At ask time: retrieve, stuff into a prompt with citations, refuse if the score is junk. Log the chunk IDs so we can debug a bad answer.
The messy part
Scanned PDFs, three versions of the same policy, and tables. We tell you in the quote if the corpus needs a cleanup week. That week is not free.
Questions
Can it search Google?
Not in the default build. Your corpus, your answers. Web search is a separate tool with its own guardrails.