You will re-chunk and re-embed your corpus several times: a better splitter, a cheaper model, a bug in the PDF extractor. A schema that treats those as exceptional turns each one into a migration with downtime. A schema keyed on content hashes turns them into a job you can run twice with no ill...
Source: [Dev.to](https://dev.to/multigrid/a-documents-table-that-survives-re-indexing-21eg)