| Metadata | Value |
|---|---|
| Status | Active |
| Version | 1.2.0 |
| Last Updated | 2026-09-10 |
| Author | Sangeetha Grantha Team |
| Document Type | Current guide |
AI supports extraction, transliteration, and retrieval in Sangeetha Grantha. Source fidelity, canonical identity, and publication remain controlled application/editorial responsibilities. The current architecture does not treat model output as a scholarly authority.
| Capability | Current path | Boundary |
|---|---|---|
| Source extraction | Python HTML/PDF strategies, canonical Pydantic output, Kotlin consumption | Deterministic/source-aware parsing is the primary path; supported strategies govern formats |
| Optional enrichment | Worker Gemini integration | Explicit opt-in and credentials; default disabled |
| Admin transliteration | Kotlin transliteration service and editor operation | Does not generate new reader text in Rasika automatically |
| Reference candidates | Normalization, aliases, candidate matching, curator resolution | Similarity is not proof of musical identity |
| Hybrid/semantic retrieval | Document/query embeddings, pgvector, profile compatibility, rank fusion | Separate indexing operation; not a conversational mobile assistant |
The worker README, ingestion guide, and search guide describe the operating paths.
Generation/enrichment and embeddings use different configuration and contracts. Keep model choice, response validation, retries, timeouts, and index compatibility explicit. Worker enrichment uses typed/configured SDK behavior; backend service calls retain their own clients/configuration.
Canonical payloads must preserve source text, script/language, section order, variants, unknown classification, and raga ordering. Retain extractor/model context where the pipeline provides it and inspect low-confidence or conflicting results in curation.
Use deterministic fixtures and controlled provider failures for code tests. Use source comparison and a dated corpus/query benchmark for extraction/retrieval quality. Historical cost, throughput, or accuracy estimates in old plans are not current measured performance or vendor pricing.
Record the model/profile, corpus scope, revision, commands, and observed results. Keep provider-backed evaluations separate from offline test totals. Quality and TRACK-108 evidence show the reporting boundary.
Conversational discovery, broader automated musicological validation, and expanded editorial automation require separate specifications and evidence. They are not implied by the existing embedding index or the presence of a validation route.
AI opportunities · Configuration · Feature map