Converting heterogeneous ingested content — documents, messages, tickets, code — into a structure that can actually be searched: chunking long documents into retrievable units, generating embeddings, and attaching metadata. Chunk boundaries matter more than model choice for downstream answer quality.