GenAI Engineer

Sai Narayana B

I build multi-agent systems, RAG pipelines, and ship open-source tools that make AI development faster. Currently at Quantiphi.

About Me

I work at the intersection of large language models and production engineering. At Quantiphi, I've built multi-agent systems that cut pipeline runtimes by 95%, shipped NL2SQL agents for enterprise platforms, and optimized inference costs by 80% through prompt restructuring and caching strategies.

Outside work, I ship open source โ€” my tool patchwork-conventions is published on PyPI and MCP Registry, and integrates with Claude Code, Cursor, and GitHub Copilot.

๐Ÿ“ Bengaluru, India๐ŸŽ“ B.Tech in Computer Science Engineering ยท SRM Institute of Science and Technology๐Ÿ“Š CGPA: 9.17๐Ÿ“… 2021 - 2025

1.5+

Years Experience

95%

Runtime Reduction

4

Certifications

Work Experience

Machine Learning Engineer (Applied GenAI)

Quantiphi ยท Bengaluru, India

Jul 2025 - Present

  • โ–นArchitected a multi-agent staffing system (TEX Agent) using Google ADK that dynamically spawns parallel agents for each required role and returns top candidate recommendations with fit rationale.
  • โ–นDeveloped a SoW ingestion pipeline processing 2,000+ project documents from Google Drive, summarizing via LLM, and indexing into Vertex AI Vector Search for RAG retrieval.
  • โ–นReduced pipeline runtime by 95% (45 min โ†’ 5 min) by replacing sequential role processing with dynamic parallel orchestration.
  • โ–นDesigned a conversational NL2SQL agent serving 50+ PMs and business leads, converting plain-English queries to BigQuery SQL via an MCP server.
  • โ–นRan implicit vs. explicit caching POC and restructured all agent system prompts, cutting daily inference cost from $4,000 to $2,500 (37% reduction).
  • โ–นSet up dynamic LLM routing across Gemini and Claude model families to pick cheaper/faster models based on task complexity, reducing inference cost across production agents.
  • โ–นCreated REST APIs to sync the internal knowledge base (LLM Wiki) with GCS, keeping 500+ documents used by production AI agents up to date.
  • โ–นDelivered a client demo where AI agents generated specs, built a multi-service backend with frontend, tested, and deployed it end-to-end using spec-driven development; presented to KeyBank stakeholders.
  • โ–นAudited and improved Codeaira (internal AI coding assistant) by benchmarking against 5 open-source tools, rewriting system prompts, and adding slash-command based skill invocation.
  • โ–นEngineered Cypher, a developer analytics tool used by 2,000+ employees that captures AI interaction logs from Copilot, Kiro, and Claude Code, scores prompt quality using LLM analysis, and calculates effort-saved metrics per developer.

Things I've Built

patchwork-conventions

Developed in 2 days. Published to PyPI and MCP Registry. Integrates with Claude Code, Cursor, and GitHub Copilot.

AST-based codebase scanner that auto-generates CONVENTIONS.md for AI coding agents. Tree-sitter analysis detecting naming conventions, import patterns, error handling, testing frameworks, and API shapes across 5 languages with confidence scores and real code examples.

PythonTree-sitterMCPPyPICLI

TEX Agent

95% runtime reduction โ€” 45 min down to 5 min via dynamic parallel orchestration.

Multi-agent staffing system using Google ADK that takes a project ID or role description, dynamically spawns parallel agents for each required role, and returns top candidate recommendations with fit rationale.

Google ADKRAGVertex AIBigQueryPython

NL2SQL Agent

Enables non-technical users to query complex datasets using plain English.

Conversational NL2SQL agent that lets PMs and business leads query project health data in plain English, converts it to BigQuery SQL via an MCP server, and returns summarized results.

MCPBigQueryNL2SQLPythonFastAPI

Cypher

Used by 2,000+ employees. Quantifies developer productivity gains from AI-assisted coding.

Developer tool used by 2,000+ employees that captures AI interaction logs from Copilot, Kiro, and Claude Code, scores prompt quality using LLM analysis, and calculates effort-saved metrics per developer.

PythonLLMAnalytics

Codeaira โ€” Internal AI Coding Assistant

Internal developer tool used across engineering teams at Quantiphi.

Audited and improved the internal AI coding assistant (CLI + VS Code extension) by benchmarking against open-source tools, rewriting system prompts, and adding slash-command based skill invocation.

TypeScriptVS Code APILLMCLINode.js

SoW Summarization Pipeline

Processes 2,000+ project documents. Powers the knowledge base used by production AI agents.

Ingestion pipeline that pulls 2,000+ project documents from Google Drive, summarizes them via LLM, and indexes into Vertex AI Vector Search for the staffing agent's RAG retrieval.

PythonLLMVertex AI Vector SearchGCSFastAPI

LLM Caching & Cost Optimization

Cut daily inference cost from $4,000 to $2,500 (37% reduction).

Ran implicit vs. explicit caching POC and restructured all agent system prompts. Set up dynamic LLM routing across Gemini and Claude model families to pick cheaper/faster models based on task complexity.

PythonVertex AIPrompt EngineeringCaching

Tech Stack

AI & ML

Multi-Agent OrchestrationLarge Language ModelsRAG PipelinesNL2SQLPrompt EngineeringPrompt CachingChain-of-ThoughtFew-Shot PromptingToken OptimizationLLM EvaluationAgentic AIMCP ServersA2A ProtocolFunction CallingModel RoutingGuardrails

Agent Frameworks & LLMs

Google ADKLangChainLangGraphClaude Agent SDKCopilot SDKOpenAI APIGemini APIAnthropic APILLM WikiGraphify

Languages & Backend

PythonSQLFastAPIStreamlitREST APIsTree-sitter

Cloud (GCP)

Vertex AIBigQueryCloud StorageCloud Run

Dev Tools

VS Code ExtensionsFAISSVertex AI Vector SearchEmbeddingsGitGitLab CI/CD

Get in Touch

I'm always open to discussing AI engineering, multi-agent systems, or interesting collaboration opportunities. Feel free to reach out.

Say Hello