An internal engineering case study · September 2026
Making cemetery knowledge searchable—and keeping it current
Documents, scanned PDFs, and meeting recordings hold valuable operational knowledge. A useful search system needs to keep up with new material, preserve its sources, and control what enters its index.
Babic Consulting commissioned Cemetery Brain's self-hosted indexing and document-processing infrastructure on its server platform, known internally as Spock. The work combined an NVIDIA T4 GPU with document preparation, transcription, scheduled publishing, and verification controls.
≈16×
Embedding throughput in a bounded benchmark
143,321
Searchable passages after the September 2 publish
281
Scanned PDFs processed in the recorded OCR batch
The challenge
Source collection was working, but updates were not reaching the searchable catalog. The index had remained at 79,410 passages while prepared material accumulated. Publishing depended on a separate workstation, creating another point where the process could stop.
What changed
We moved GPU indexing onto the server that holds the data and removed the workstation dependency. We verified GPU execution, measured throughput against the CPU baseline, and checked sustained-load behavior before completing the catch-up.
The pipeline publishes a new catalog only after integrity and health checks pass. It preserves source provenance and includes rollback controls. Documents pass through defined approval and privacy-processing stages before becoming searchable.
For meetings, the system prefers an existing Teams transcript and uses local speech recognition when a recording has no transcript. Scanned documents receive OCR. Each stage can find unfinished work and run again without blindly duplicating completed output.
Measured results
- A 512-passage benchmark measured 34.04 passages per second on the NVIDIA T4 versus 2.133 on the CPU—about 16× higher embedding throughput. This measures search-index preparation, not chatbot response speed.
- The catch-up run embedded 55,627 passages. A subsequent unattended publish brought the catalog to 143,321 passages on September 2, 2026.
- The recorded OCR batch processed 281 scanned PDFs with zero processing errors. Completion does not establish perfect text recognition.
- Local transcription converted 69 previously untranscribed recordings to text. Five silent or no-audio recordings were explicitly classified, with zero unhandled failures in the completed backlog.
The result is an operating foundation for keeping approved cemetery knowledge searchable on business-controlled infrastructure. The next evaluation focuses on whether generated answers support real staff questions with useful coverage and grounded sources. These infrastructure results do not establish staff time savings or answer accuracy.
Put your operational knowledge to work
Have valuable knowledge spread across documents, scans, and recordings? We can assess your sources, infrastructure, and workflows, then define a practical first use case with measurable acceptance criteria.
Talk to Andrew