Sangeetha-Grantha

Metadata Value
Status Active
Version 1.1.0
Last Updated 2026-09-10
Author Sangeetha Grantha Team
Document Type Evidence record

Bulk Import Fixes - Implementation Plan


[!NOTE] Historical evidence: results, counts, commands, and observations below belong to the original work described here. The editorial update date is not a new test or corpus verification. For present behavior, use current quality checks.


Executive Summary

This document consolidates findings from three comprehensive code reviews (Claude, Goose, Codex) and provides a prioritized implementation plan to address critical issues, gaps, and improvements in the bulk import system.

Review Summary

Overall Assessment:

Key Findings:


Issue Categorization

Critical (Must Fix - Blocking Production)

  1. Manifest Ingest Failure Handling (All 3 reviews)
    • Batch not marked FAILED when manifest ingest fails with zero tasks
    • Violates clarified requirements (2026-01)
    • Track: TRACK-010
  2. Task Stuck Detection Race Condition (Codex)
    • Tasks marked RUNNING at claim time, but watchdog may mark as RETRYABLE before execution
    • Risk of double-processing and duplicate side effects
    • Track: TRACK-010
  3. File Upload Security Vulnerabilities (Codex, Goose)
    • Path traversal risk (no filename sanitization)
    • No file size limits (OOM risk)
    • Null filename handling
    • Track: TRACK-010
  4. Quality Scoring System Missing (All 3 reviews)
    • Strategy specifies quality tiers (EXCELLENT, GOOD, FAIR, POOR)
    • No calculation or persistence of quality scores
    • Blocks review prioritization and automation
    • Track: TRACK-011

High Priority (Should Fix - Quality & Completeness)

  1. Review Workflow APIs Incomplete (Claude, Goose)
    • Missing bulk-review and auto-approve queue endpoints
    • Limited batch-scale moderation workflows
    • Track: TRACK-012
  2. Auto-Approval Logic Incomplete (All 3 reviews)
    • Hardcoded heuristics instead of configurable rules
    • Missing quality tier integration
    • Track: TRACK-012
  3. Entity Resolution Cache Issues (Goose, Codex)
    • In-memory only, no database persistence
    • Cache invalidation missing when new entities created
    • Multi-node deployment divergence risk
    • Track: TRACK-013
  4. Deduplication Service Incomplete (All 3 reviews)
    • Missing intra-batch deduplication during processing
    • O(N^2) performance for large batches
    • Track: TRACK-013

Medium Priority (Performance & Scalability)

  1. Stage Completion Checks O(N) per Task (Claude, Codex)
    • checkAndTriggerNextStage loads all tasks on every completion
    • ~1,200 unnecessary DB queries per batch
    • Track: TRACK-013
  2. Rate Limiting Too Conservative (Claude, Codex)
    • Defaults (12/min per domain, 50/min global) vs strategy (120/min)
    • 5-10x slower than strategy estimates
    • Track: TRACK-013
  3. CSV Validation Deferred (Codex)
    • Validation happens at manifest ingest, not upload
    • Invalid CSVs fail minutes later instead of fast-failing
    • Track: TRACK-013
  4. CSV Parsing Issues (Codex)
    • Platform default charset (diacritic handling risk)
    • File readers not closed (file descriptor leaks)
    • Track: TRACK-010

Low Priority (Code Quality & Polish)

  1. Normalization Bugs (Codex)
    • Honorific removal regex broken ("\b" is backspace, not word boundary)
    • Normalized lookup maps drop collisions
    • Track: TRACK-013
  2. Memory Leak in Rate Limiter (Claude)
    • perDomainWindows map grows unbounded
    • Track: TRACK-013
  3. Testing Gaps (All 3 reviews)
    • No unit tests for normalization/resolution logic
    • No integration tests for full pipeline
    • Track: TRACK-014
  4. Frontend Improvements (Goose)
    • Missing batch filter in review queue
    • No quality tier filtering
    • Track: TRACK-012

Implementation Tracks

TRACK-010: Critical Fixes & Security Hardening

Priority: CRITICAL
Estimated Effort: 2-3 days
Status: Proposed

Scope:

Dependencies: None
Blocks: Production deployment


TRACK-011: Quality Scoring System

Priority: HIGH
Estimated Effort: 3-4 days
Status: Proposed

Scope:

Dependencies: None
Blocks: Review prioritization, auto-approval automation


TRACK-012: Review Workflow Completion

Priority: HIGH
Estimated Effort: 4-5 days
Status: Proposed

Scope:

Dependencies: TRACK-011 (quality scoring)
Blocks: Full automation, batch-scale moderation


TRACK-013: Performance & Scalability Improvements

Priority: MEDIUM
Estimated Effort: 5-7 days
Status: Proposed

Scope:

Dependencies: None
Blocks: Scale beyond 5,000 krithis


TRACK-014: Testing & Quality Assurance

Priority: MEDIUM
Estimated Effort: 5-7 days
Status: Proposed

Scope:

Dependencies: TRACK-010, TRACK-011, TRACK-012, TRACK-013
Blocks: Confidence in production stability


Implementation Timeline

Phase 1: Critical Fixes (Week 1)

Phase 2: Quality & Review (Weeks 2-3)

Phase 3: Performance & Testing (Weeks 4-5)


Success Criteria

TRACK-010 (Critical Fixes)

TRACK-011 (Quality Scoring)

TRACK-012 (Review Workflow)

TRACK-013 (Performance)

TRACK-014 (Testing)


Risk Assessment

High Risk

Medium Risk

Low Risk


Dependencies & Blockers

Critical Path

  1. TRACK-010 → Must complete before production deployment
  2. TRACK-011 → Blocks TRACK-012 (review workflow needs quality scores)
  3. TRACK-012 → Depends on TRACK-011
  4. TRACK-013 → Can proceed in parallel with TRACK-011/012
  5. TRACK-014 → Should follow TRACK-010-013 completion

External Dependencies


Notes

Clarified Requirements (2026-01)

The implementation correctly follows most clarified requirements:

Architecture Decisions

Performance Considerations


References


Implementation Plan Created: 2026-01-23
Next Review: After TRACK-010 completion


Section index · Documentation home · Feature status