| Metadata | Value |
|---|---|
| Status | Actioned (see resolution notes) |
| Version | 1.1.0 |
| Last Updated | 2026-09-10 |
| Author | Principal Data & AI Engineering review (for Seshadri) |
| Document Type | Design reference |
| Baseline reviewed | current-versions.md v1.1.0 (2026-03-10) |
| Review window | 2026-03-10 → 2026-06-06 (~3 months) |
| Resolution | F1–F3 addressed by TRACK-107 and TRACK-124. F7 by TRACKs 120–124. F6 semantic search shipped as TRACK-108 (2026-09). F4–F5 remain opportunities. |
[!NOTE] Design/reference material: this page may include proposals or earlier implementation assumptions. Use current feature map for implemented behavior and current operating steps.
A dependency and AI-ecosystem review after a ~3 month gap. The goal is a short list of changes that are practical, relevant, and worth your time — and an equally explicit list of changes that are noise you should ignore. Findings are severity-rated. Nothing here is recommended on novelty alone.
The codebase itself has aged gracefully — most of your library versions are still current or one minor release behind, which is fine. The risk is not in your build files; it is in your AI layer. Two of the three Gemini models this project depends on have reached or passed Google’s retirement dates, and the Python AI SDK you ship in the extraction worker was deprecated last year with its migration deadline already lapsed. These are not “upgrade when convenient” items — they are correctness-and-availability items that will cause hard failures (HTTP 404 on the model endpoint, no security patches on the SDK) if left.
The good news is that the same churn that created these breaks also opened three genuinely useful, low-hype opportunities that map directly onto features already named in your own 09-ai/integration-opportunities.md: Batch Mode (50% cheaper bulk processing), first-class JSON-Schema structured output (kills your hand-rolled prompt-and-pray JSON parsing), and a GA embeddings model that finally makes the “semantic search” opportunity buildable without standing up a separate ML stack.
| # | Finding | Severity | Area |
|---|---|---|---|
| F1 | Gemini 2.0 Flash (your primary model) reached retirement on 1 June 2026 | Blocker | AI |
| F2 | Gemini 1.5 Pro (musicological validation) is deprecated | Critical | AI |
| F3 | google-generativeai 0.8.6 Python SDK deprecated; migration deadline already passed |
Critical | AI / Python |
| F4 | No structured-output (JSON Schema) usage — extraction relies on prompt-coerced JSON | Major | AI |
| F5 | Bulk transliteration/scraping runs at full synchronous price; Batch Mode would halve it | Major | AI / Cost |
| F6 | “Semantic search” opportunity remains unbuilt; embeddings are now GA and trivial to adopt | Major | AI / Feature |
| F7 | Backend one minor behind (Kotlin 2.4, Ktor 3.5 available) | Minor | Backend |
| F8 | Frontend already on the new @google/genai SDK — no action, just verify model strings |
Observation | Frontend |
| F9 | Gemma 4 12B (open-weight) — not a fit for the core path; one eval-gated fine-tune angle worth watching | Observation | AI (see §5) |
The model lineup moved a full generation in three months. Here is the lineage relevant to you:
Gemini 1.5 Pro → deprecated (your validation model)
Gemini 2.0 Flash → RETIRED 1 Jun 2026 (your primary model)
↓ replaced by
Gemini 2.5 Flash → stable, but itself slated to shut down 16 Oct 2026
↓ replaced by
Gemini 3.5 Flash → GA since 19 May 2026, no shutdown date announced
gemini-selection-rationale.md selects Gemini 2.0 Flash as the workhorse for transliteration, scraping, and metadata normalisation, and Gemini 1.5 Pro for musicological validation. Google retired gemini-2.0-flash-001 (and -flash-lite-001) on 1 June 2026. Any code path still pinning that model string is now calling a dead endpoint or being silently redirected — neither is acceptable in a system of record.
Recommendation — skip a generation, do not chase the treadmill. Migrating to Gemini 2.5 Flash buys you only ~4.5 months before its own 16 Oct 2026 shutdown forces a second migration. Go straight to Gemini 3.5 Flash (GA, no announced retirement). It is reported to be ~4× faster on output tokens than the 2.x line, beats the previous 3.1 Pro tier on hard benchmarks, and lands around $1.50 / million input tokens. For your batch-heavy, extraction-and-transliteration workload that is the right tier on speed, cost, and longevity simultaneously.
This is a config-and-eval change, not a re-architecture — but treat the eval seriously: re-run your Indic transliteration golden set (the La/Lla/Zha and Grantha-vs-Tamil distinctions you flagged) before flipping production. Newer ≠ automatically better on your niche; verify, don’t assume.
The “musicological validation” task (checking Arohanam/Avarohanam against lyrics) is the one place you deliberately chose a heavier reasoning model. With 1.5 Pro deprecated, consolidate onto Gemini 3.5 Flash here too and A/B it against the (soon-to-roll-out) 3.5 Pro tier only if Flash’s reasoning proves insufficient on your validation set. In practice the 3.x Flash reasoning is now in the range that previously required a Pro model, so you may be able to collapse two model tiers into one — simpler billing, one prompt library, one place to monitor drift.
google-generativeai SDK is deprecated [CRITICAL]Your extraction worker (tools/krithi-extract-enrich-worker) ships google-generativeai==0.8.6. That package was deprecated on 30 November 2025, and the Gemini-API migration deadline (31 August 2025) has already passed. It does not receive new features (Batch Mode, Live API, the newest structured-output surface) and is on borrowed time for fixes.
Migrate to the unified google-genai SDK (currently ~1.62.x). This is not cosmetic — it is the precondition for F4–F6 below. The migration is mechanical but real: the old SDK configured a global client implicitly (genai.configure(api_key=...)); the new one is client-first (client = genai.Client(...) then client.models.generate_content(...)). One client, both backends (Gemini Developer API and Vertex AI) behind the same interface — which matters because gemini-selection-rationale.md already anticipates a Vertex AI / unified-billing path on GCP. The new SDK lets you switch Developer-API ↔ Vertex with a constructor argument instead of a rewrite.
Frontend note (F8): the admin web already depends on
@google/genai 1.34.0— that is the new unified JS SDK, so the frontend is on the right side of this split. The only action there is to confirm no hard-codedgemini-2.0-flash/gemini-1.5-prostrings remain.
Your integration docs show transliteration/extraction prompts that end with “Provide ONLY the … text, no explanations” — i.e. you are coercing format through prompt wording and then parsing optimistically. Google has since added full JSON Schema support across all current Gemini models, and the unified SDKs accept a Pydantic model (Python) or Zod schema (JS) directly as response_schema. You already run Pydantic 2.12 in the extraction worker, so this is nearly free leverage:
response_schema=KrithiExtraction Pydantic model.imported_krithis.parsed_payload, less defensive parsing, fewer reviewer rejections caused by format drift rather than content.This is the single highest effort-to-payoff item in the report.
The Gemini Batch API processes large asynchronous jobs (results within 24h) at 50% of synchronous price, and Batch requests can carry the same response_schema from F4. Your scraping/transliteration backfills (the “thousands of compositions” ambition in the rationale doc) are the textbook fit: they are not latency-sensitive, they are embarrassingly parallel, and they run as backfills not user-interactive calls.
Keep synchronous calls only where a human is waiting (the “Generate Variants” button in KrithiEditor.tsx). Route everything that originates from the import pipeline or a scheduled backfill through Batch. On Gemini 3.5 Flash pricing, halving the bulk spend is a material, recurring saving with no quality trade-off — the same model, just scheduled.
Opportunity #5 in your own integration-opportunities.md (“semantic search beyond keyword matching”) has been parked. The blocker — needing an embeddings model and a vector store — is now much smaller:
gemini-embedding-001 is generally available, multilingual across 100+ languages (covers your Dravidian + Sanskrit corpus), with Matryoshka Representation Learning: emit 3072-dim vectors and truncate to 1536 or 768 with minimal quality loss. Start at 768 dims to keep your index small and cheap; raise only if recall demands it. (Google has also announced Gemini Embedding 2, a natively multimodal successor — note it, but build on the GA 001 model today; don’t gate a shippable feature on a just-announced one.)pgvector keeps embeddings beside your relational data — no separate vector store, no new operational surface, consistent with the “open-source, minimise integration debt” posture. A krithi_embedding table (krithi_id, section_id, vector, model_version) indexed with HNSW is enough to ship “find similar krithis / search by meaning” over the existing catalogue.This is the one new feature (vs. maintenance) I’d actively pull forward, because it converts work you’ve already documented into a user-visible capability with a small, well-understood footprint.
Most of the stack is current. Listing only what moved and whether it’s worth doing.
| Library | You have | Now available | Action | Notes |
|---|---|---|---|---|
| Kotlin | 2.3.0 | 2.4.0 (Jun 2026); 2.3.20 also out | Optional minor | Routine. Adopt on your normal cadence; no forcing function. |
| Ktor | 3.4.0 | 3.5.0 (15 May 2026) | Optional minor | 3.4 added OpenAPI generation, Zstd compression, structured-concurrency request lifecycle. Worth reading the 3.5 changelog for the OpenAPI gen — you maintain openapi/sangita-grantha.openapi.yaml by hand today. |
| Exposed | 1.0.0 | 1.0.x patches | Hold / patch only | You’re already on the landmark 1.0 (R2DBC support, the major rewrite). Good place to be. Just take patch releases. |
| Koin, Coroutines, HikariCP, Logback, Caffeine, PostgreSQL driver | current-ish | minor/patch bumps | Batch as housekeeping | None urgent. Roll up in one dependency-bump PR. |
Verdict: no backend item rises above Minor. Do a single housekeeping bump PR when convenient; don’t interrupt feature work for it. Exposed 1.0 was the only thing that could have been disruptive and you already absorbed it.
| Library | You have | Status | Action |
|---|---|---|---|
| React | 19.2.4 | 19.x still current | Patch only |
| Vite | 7.3.1 | 7.x current | Patch only |
| Tailwind CSS | 4.2.1 | 4.x current | Patch only |
| TanStack Query | 5.90.x | v5 current | Patch only |
| React Router | 7.13.x | v7 current | Patch only |
@google/genai |
1.34.0 | already the new unified SDK | Verify model strings only |
Verdict: the frontend is in good shape and notably already past the SDK split that bites the Python worker. No structural work. Only confirm there are no retired model strings hard-coded in the “Generate Variants” path.
| Library | You have | Now | Action | Severity |
|---|---|---|---|---|
google-generativeai |
0.8.6 | deprecated → google-genai ~1.62.x |
Migrate | Critical (F3) |
| Pydantic | 2.12.5 | 2.12.x stable (2.13 in beta) | Hold; leverage for F4 response_schema |
— |
| PyMuPDF | 1.27.1 | 1.27.2.x patches | Patch | Minor |
| pdfplumber, psycopg, RapidFuzz, HTTPX, structlog, etc. | current | minor/patch | Housekeeping | Minor |
| Ruff / mypy / pytest | current | minor | Housekeeping | Minor |
Verdict: one Critical (the SDK migration, F3), everything else is housekeeping. Bundle the housekeeping with the SDK migration since you’ll be in pyproject.toml/uv.lock anyway.
Ordered by urgency × payoff, with rough effort. This is a maintenance-plus-one-feature plan, not a rewrite.
Now / this sprint (availability & correctness — non-negotiable):
google-genai. ~0.5–1 day. Unblocks everything else.gemini-3.5-flash behind a single config constant; re-run the Indic transliteration + raga-validation golden set before promoting. ~1 day incl. eval. Track it as a Conductor TRACK-XXX per your 09-ai/README.md convention, and update current-versions.md + the three sync targets in CLAUDE.md.Next (cost & quality, high payoff, low risk):
response_schema. ~1–2 days. Biggest reliability win per unit effort.Then (the one new capability worth pulling forward):
gemini-embedding-001 + pgvector. New krithi_embedding table, offline backfill job, similarity endpoint. ~1 week for a usable v1 at 768 dims. Register as its own Conductor track; this is a feature, not maintenance.Housekeeping (any time, one PR each):
@google/genai calls.Requested follow-up: can an open-weight Gemma 4 12B model take over any of the AI work — including function calling and multimodal — and where is that a good idea versus a distraction? Verdict up front, evidence after.
Gemma 4 12B shipped 3 June 2026 and is a genuinely strong release on paper: Apache 2.0 licence (no more custom “Gemma Terms” carve-outs — you can fine-tune, redistribute, and run it commercially with only attribution obligations), an encoder-free unified multimodal architecture (text/image/audio/video in, text out), a 256K context window, 140+ languages, a custom tool-call protocol for function calling, and a memory footprint small enough to run on ~16 GB VRAM or unified memory at 4-bit. It beats Gemma 3 27B on standard benchmarks and nearly matches the 2×-larger 26B MoE. It is available hosted on Vertex AI Model Garden / Cloud Run or fully self-hosted (vLLM for GPU throughput, llama.cpp/GGUF for Apple Silicon).
None of that is in dispute. The question is whether it fits this product, and the honest answer is: mostly no for the core path, with two narrow, eval-gated exceptions. [Severity: Observation — opportunity, not obligation]
1. The core task is Indic transliteration fidelity, and that is exactly where a general 12B open model is weakest. Your entire Gemini rationale (gemini-selection-rationale.md) hinges on niche correctness — preserving La/Lla/Zha distinctions and not bleeding Tamil into Grantha. “140+ languages” is a statement about breadth, not depth on Dravidian-script edge cases. A 12B model that is excellent at GPQA and DocVQA tells you nothing about whether it renders a Sanskritised Telugu kriti correctly. Assume it is worse than Gemini 3.5 Flash here until a golden-set eval proves otherwise — do not take the benchmark-leader framing at face value for your domain.
2. The cost case for self-hosting does not close at your volume. This is the part the marketing (“60–80% cheaper at scale”, “runs on a laptop”) obscures. The savings are real only when a GPU stays busy. Gemini 3.5 Flash is ~$1.50/M input tokens, halved again via Batch Mode (F5) — you are already on a near-floor price for bursty work. Self-hosting means a 16–24 GB GPU (L4/A10-class) running ~$0.5–1.0/hr; left on, that is ~$400–700/month before you process a single token. Your profile is a curated archive of thousands of compositions processed largely once, plus occasional interactive edits — bursty and low-duty-cycle. To break even against managed Flash-Batch you would need sustained, high-volume throughput you do not have. For your FinOps reality, the managed API is the cheaper option, not the more expensive one. Self-hosting to “save money” here would increase total cost once you price in the GPU and the engineer-hours.
3. It adds operational surface that contradicts your stated posture. Self-hosting means GPU provisioning, a serving stack (vLLM), quantisation management, drift/throughput monitoring, and scaling — net-new ops for a small team on a system of record that explicitly values determinism and auditability. Every hour spent on model-serving infra is an hour not spent on the catalogue. The managed path keeps your AI surface a stateless API call behind a config constant.
4. Function calling solves a problem you don’t have. Gemma 4’s tool-calling is competent (the widely-cited ~17/20 chain-completion figure was measured on the larger 31B dense variant, not the 12B — be skeptical of transferring it down). But your AI surface is bounded extraction → transliteration → validation → human review. That is not an agent orchestrating tools; it is structured generation. Native JSON-Schema structured output (F4) is the right primitive for your need, and Gemini already does it well. Adopting Gemma for its function-calling would be buying a capability your architecture deliberately doesn’t use.
5. Multimodal is a narrow, real, future adjacency — not a reason to switch now. Two honest possibilities: (a) Vision could OCR scanned/image-only PDFs where PyMuPDF/pdfplumber return nothing — a genuine fallback for poor manuscript scans, but a niche tail, not the main pipeline. (b) Audio/ASR is interesting if the product ever ingests recorded renditions to derive notation — a speculative roadmap item, not current scope. Neither justifies displacing Gemini today.
Two specific, defensible bets — both eval-gated, neither urgent:
For embeddings (F6), Gemma is not the tool — it is a generative model, not an embedding model. Keep gemini-embedding-001. Don’t conflate the two because they share a brand.
Do not move core transliteration, validation, or extraction off Gemini 3.5 Flash onto a self-hosted Gemma 4 12B. The fidelity risk is real, the cost case is inverted at your volume, the ops burden contradicts your posture, and the function-calling/agentic capabilities answer a question your architecture doesn’t ask. The only strategically interesting use is a future fine-tuned, domain-owned transliteration model — and that is a deliberate, eval-gated investment to make after you have a labelled corpus, not a swap to do now. Run experiment #1 to quantify the gap; treat everything else as a watching brief.
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Production extraction fails on retired 2.0 Flash endpoint | High (already past retirement) | High | F1 — repoint model string immediately |
| Model swap degrades Indic transliteration fidelity on edge scripts | Medium | High | Golden-set eval before promotion; keep prompts schema-constrained (F4) for stability |
google-genai migration introduces auth/client regressions |
Low | Medium | Migrate in a branch; the worker already isolates the Gemini client |
| Re-migration churn (3.5 Flash later retired) | Low | Low | Single config constant for model string; scheduled re-eval, not reactive chasing |
| pgvector adds operational load | Low | Low | Stays inside existing Postgres 18; HNSW index; no new datastore |
| Batch Mode 24h latency leaks into user-interactive paths | Low | Medium | Restrict Batch to import/backfill; keep “Generate Variants” synchronous |
| Premature self-hosting of Gemma 4 12B increases cost & ops for no gain | Medium (tempting on paper) | Medium | Section D — managed Gemini is cheaper at your volume; gate any Gemma move on golden-set eval + a real fine-tune corpus |
You did not fall behind on libraries — your build files are fine and the few minors can wait. You fell behind on AI platform lifecycle, which moves faster than dependency semver and breaks harder. Spend the next sprint on F1–F3 (availability), bank the cheap reliability/cost wins in F4–F5, and treat F6 as the single feature investment that turns documented intent into something users can feel. Ignore the agentic/multimodal hype entirely for this product — a curated system of record wants determinism and auditability, and that is exactly what the boring, schema-constrained, batch-processed path gives you.
AI ecosystem:
Libraries:
Gemma 4: