Building scalable backend systems and data pipelines for real-world applications — from event-driven Kafka pipelines processing millions of records/day to RAG systems serving answers in <200ms and REST APIs with 30% faster response times.
Designed and built a two-service RAG system — async ingestion pipeline (chunking, embedding, 10K+ docs indexed) and FastAPI query service with cross-encoder re-ranking. Sub-200ms latency, 94% retrieval accuracy, zero hallucinations, and full LLM cost observability via MLflow.
Enterprise teams spend hours searching unstructured internal documents — PDFs, wikis, reports. Keyword search returns noise with no context. Generic LLMs hallucinate without domain grounding. The gap: fast, accurate, cited answers from private knowledge bases.
A two-layer RAG system: an ingestion pipeline that chunks, embeds, and stores documents in a Vector DB, and a query pipeline that retrieves semantically relevant context and passes it to the LLM — ensuring grounded, cited responses with full observability.
Designed and built the full system end-to-end — ingestion pipeline, embedding strategy, retrieval logic, FastAPI service, and MLflow observability layer. Made all core architecture decisions including chunking strategy, vector DB selection, and LLM orchestration approach.
Sub-200ms query latency. Zero in-domain hallucinations — every answer cites its source chunk. First-ever LLM cost dashboard giving visibility into token usage and cost-per-query. Stateless architecture ready for horizontal scaling.
Documents are split using semantic chunking to preserve paragraph coherence. Each chunk is embedded and stored with metadata: source file, page number, timestamp, and chunk ID for precise citations at query time.
Query-time embedding compared via cosine similarity. Top-K=20 retrieved, re-ranked by cross-encoder to top-5. Prompt template enforces the model only answers from provided context — eliminating hallucination.
Stateless async FastAPI with UUID traceability per request. Async endpoints prevent blocking on LLM calls. Middleware handles auth, rate limiting, and request timing — feeding every metric into the observability layer.
Four deliberate architecture decisions that shaped the system.
Expensive GPU compute, long cycles, retraining every time docs change. Not viable for a dynamic corpus.
No retraining needed — add a doc, re-embed it. Real-time updates, lower cost, same accuracy.
Fast but splits mid-sentence, breaking semantic coherence. Hurts retrieval quality.
94% vs ~78% retrieval accuracy in internal benchmarks.
Blocks FastAPI worker threads during LLM latency. Collapses throughput under load.
10× higher concurrent request handling.
Ingestion and query have different scaling needs. Can't scale independently.
Query scales independently during peak. Each service has its own failure domain.
Noisy top-K results — semantically similar but contextually irrelevant chunks slipping into the retrieved set, confusing the LLM.
Added cross-encoder re-ranking on top-20 before trimming to top-5. Accuracy: ~81% → 94%.LLM API tail latency hit 1.8s vs 600ms median — making SLA commitments impossible.
Streaming responses start rendering at first token (~180ms). Parallel async retrieval cut pre-LLM latency to <20ms.Updated documents kept old embeddings in Vector DB — users got answers citing superseded versions.
SHA-256 fingerprinting detects changes and triggers selective re-embedding — only modified chunks reprocessed.No real-time cost tracking — no way to know which queries were expensive or if prompt changes were cost-effective.
Every LLM call logged with token counts, cost estimate, latency → SQL + MLflow dashboard. Enabled 30% prompt optimisation.LLM calls take 500ms–2s. Sync blocks the worker thread. Async + await serves other requests while waiting — 10× throughput improvement under concurrent load.
"Answer ONLY from context. Cite source IDs." — forces the LLM to stay within retrieved context. If the answer isn't there, it says so rather than inventing one.
Designed and deployed an event-driven Kafka + Airflow distributed pipeline processing millions of log events/day with fault-isolated DAG stages. Optimised PostgreSQL via composite indexing and execution plan tuning — 35% latency reduction, 40% manual ETL ops eliminated.
Log data from multiple services arrived unstructured and at high volume. Manual processing caused delays, data loss, and 40% ops overhead — no scalable way to query historical events.
Event-driven pipeline: Kafka topics per log source → Airflow DAGs for fault-isolated transformation stages → PostgreSQL with composite indexes for analytical queries.
Chose Kafka over direct DB writes for decoupling — producers don't block on consumer lag. Dead-letter queue catches failed events without pipeline stalls. Consumer-lag monitoring via custom Airflow sensors.
35% query latency reduction via schema redesign and EXPLAIN ANALYZE tuning. 40% manual ops eliminated through automated DAG orchestration. Millions of events/day processed reliably.
Built end-to-end ML engineering pipeline — Pandas feature engineering → Scikit-learn classification → cross-validation (~87% precision, F1-scored) → Joblib serialisation to deploy-ready .pkl artifact. Sub-50ms inference; wrappable in FastAPI for production serving.
Business needed to identify customers at risk of churning before they left. No existing ML pipeline — raw CSV data, no feature engineering, no model, no deployment path.
Pandas ETL for feature engineering → Scikit-learn classifier with cross-validation → Joblib serialisation to .pkl → FastAPI-wrappable inference endpoint. Full retrain pipeline included.
Chose Scikit-learn over deep learning — tabular data, interpretability required, fast inference needed. Decoupled preprocessing from training so retraining only rebuilds affected stages, not full pipeline.
~87% precision on held-out test data, F1-scored to handle class imbalance. <50ms inference latency. Deploy-ready .pkl artifact wrappable in any Python REST framework in under 10 lines.
Built four-module data pipeline system — Pandas ETL → geospatial feasibility scoring (14 MA counties) → CO₂ emissions modelling → real-time delivery simulator with live Plotly KPI computation. Fully deployable to Streamlit Cloud.
Amazon Prime Air feasibility data was scattered across CSV files with no unified analysis layer. Needed a system to score delivery viability, compare emissions, and simulate delivery times across geographies.
Pandas ETL pipelines feed four independent modules: geospatial scoring engine, product compatibility filter, CO₂ emissions calculator, and real-time delivery simulator — all rendered via Plotly + Streamlit.
Chose modular Pandas pipelines over a monolithic script — each analysis module is independently composable. Adding a new delivery method or county requires one CSV file swap, no code changes.
14 Massachusetts counties scored and visualised. 3 delivery methods compared on emissions and time. Real-time KPI computation with zero backend latency. Streamlit Cloud deployable.
Software & Data Engineer
📍 Boston, MA
she / her
I'm a software engineer with a focus on backend systems, distributed data pipelines, and cloud-native architectures. I've built systems that process millions of records/day, reduced PostgreSQL query latency by 35%, cut manual data ops by 40%, and served REST APIs with 30% faster response times.
Currently finishing my M.S. in Business Analytics at UMass Amherst (graduating May 2025), I bridge the gap between engineering and data-driven decision making. I'm actively exploring opportunities in software engineering, backend, data engineering, and ML engineering.
Enterprise teams spend hours searching unstructured internal documents — PDFs, wikis, reports. Keyword search returns noise with no context. Generic LLMs hallucinate without domain grounding. The gap: fast, accurate, cited answers from private knowledge bases.
A two-layer RAG system: an ingestion pipeline that chunks, embeds, and stores documents in a Vector DB, and a query pipeline that retrieves semantically relevant context and passes it to the LLM — ensuring grounded, cited responses with full observability.
Designed and built the full system end-to-end — ingestion pipeline, embedding strategy, retrieval logic, FastAPI service, and MLflow observability layer. Made all core architecture decisions including chunking strategy, vector DB selection, and LLM orchestration approach.
Sub-200ms query latency. Zero in-domain hallucinations — every answer cites its source chunk. First-ever LLM cost dashboard giving visibility into token usage and cost-per-query. Stateless architecture ready for horizontal scaling.
Documents are split using semantic chunking (not fixed-size) to preserve paragraph coherence. Each chunk is embedded using a sentence-transformer model and stored in the Vector DB with metadata: source file, page number, timestamp, and chunk ID. This enables precise citations at query time.
Query-time embedding is computed and compared against stored vectors using cosine similarity. Top-K=5 chunks are retrieved, re-ranked by relevance score, then passed as context to the LLM. The prompt template enforces that the model only answers from the provided context — eliminating hallucination on in-domain questions.
Stateless FastAPI handles all query routing. Async endpoints prevent blocking on LLM calls. Each request gets a UUID for full traceability through the audit log. Middleware handles auth, rate limiting, and request timing — feeding every metric into the MLflow + SQL observability layer.
Every major architecture decision involved a deliberate tradeoff. Here are the four choices that most shaped the system.
Would require expensive GPU compute, long training cycles, and retraining every time docs change. Not viable for a dynamic document corpus.
No retraining needed — add a doc, re-embed it. Real-time updates, lower cost, same accuracy on in-domain questions.
Fast and simple, but splits mid-sentence, breaking semantic coherence. Retrieved chunks lack context boundaries — hurts retrieval quality.
Preserves natural language units. 94% retrieval accuracy vs ~78% with fixed chunking in internal benchmarks. Worth the extra preprocessing cost.
Simplest to implement but blocks FastAPI worker threads during LLM latency (500ms–2s). Under load, this collapses throughput.
10× higher concurrent request handling. FastAPI workers free immediately while awaiting LLM response. No degradation under moderate load.
Easier to deploy initially, but ingestion and query workloads have completely different scaling needs — you don't want to scale the whole system when only queries spike.
Query service scales independently during peak usage. Ingestion runs as a background batch job. Each service has its own failure domain.
With more documents, the Vector DB returned increasingly noisy top-K results — semantically similar but contextually irrelevant chunks were slipping into the retrieved set, confusing the LLM and producing vague answers.
Added a re-ranking layer using cross-encoder scoring on the top-20 candidates before trimming to top-5. Retrieval accuracy jumped from ~81% to 94%.The LLM API had high tail latency — median responses came back in 600ms, but the 99th percentile hit 1.8s. This made SLA commitments impossible and user experience inconsistent.
Implemented streaming responses so the UI starts rendering at first token (~180ms). Parallel async retrieval reduced pre-LLM latency to <20ms. Added a 2s timeout with graceful fallback messaging.Documents were updated frequently, but the Vector DB retained old embeddings. Users received answers citing superseded versions of internal policies — a serious trust problem in an enterprise context.
Built a document fingerprinting system (SHA-256 hash on content) that detects changes and triggers selective re-embedding — only modified chunks are re-processed, keeping ingestion costs low.Without real-time cost tracking, there was no way to know which query patterns were expensive, which users were over-consuming, or whether prompt engineering changes were cost-effective.
Instrumented every LLM call with token counts (prompt + completion), cost estimate, and latency — all logged to SQL and surfaced in the MLflow dashboard. Cost-per-query visibility enabled a 30% prompt optimization.LLM API calls take 500ms–2s. Synchronous calls block the entire FastAPI worker thread during that wait. Async + await lets the worker serve other requests while waiting, giving 10× throughput improvement under concurrent load with zero extra infrastructure.
The key to zero hallucinations: "Answer ONLY from context. Cite source IDs." This single instruction in the prompt template forces the LLM to stay within retrieved context — if the answer isn't there, it says so rather than inventing one.
I'm targeting Backend Engineer, Data Engineer, and ML Engineer roles. Here's the quick picture — what I build, how I work, and the impact I've shipped.