class Engineer(Atmika): stack = ["Python","AWS","Kafka"] ms = "Business Analytics, UMass" passion = "distributed systems" status = "open_to_work = True" def build(self): return "scalable pipelines"
Open to Full-Time Roles

ATMIKA MANOJ
PAREY

Software Engineer  |  Backend  |  Data Engineering  |  ML Systems

Building scalable backend systems and data pipelines for real-world applications — from event-driven Kafka pipelines processing millions of records/day to RAG systems serving answers in <200ms and REST APIs with 30% faster response times.

3+ Yrs Industry
40% Ops Automated
35% Query Speedup
01 AI · Backend · RAG Systems

Enterprise GenAI
Control Tower

Designed and built a two-service RAG system — async ingestion pipeline (chunking, embedding, 10K+ docs indexed) and FastAPI query service with cross-encoder re-ranking. Sub-200ms latency, 94% retrieval accuracy, zero hallucinations, and full LLM cost observability via MLflow.

<200msQuery latency
94%Retrieval accuracy
0Hallucinations
PythonLangChainVector DBFastAPILLM APIsMLflow
GitHub
System Architecture
PDF / Docs
→
Chunker
→
Embedder
→
Vector DB
User Query
→
FastAPI
→
Retriever
→
LLM
→
Response + Source
MLflow
+
SQL Analytics
→
Cost Dashboard
Async FastAPI · cross-encoder re-ranking · token cost observability
<200msEnd-to-end query latency
10K+Documents indexed
94%Retrieval accuracy
0In-domain hallucinations
3Pipeline stages
The Problem

Enterprise teams spend hours searching unstructured internal documents — PDFs, wikis, reports. Keyword search returns noise with no context. Generic LLMs hallucinate without domain grounding. The gap: fast, accurate, cited answers from private knowledge bases.

The Solution

A two-layer RAG system: an ingestion pipeline that chunks, embeds, and stores documents in a Vector DB, and a query pipeline that retrieves semantically relevant context and passes it to the LLM — ensuring grounded, cited responses with full observability.

My Role

Designed and built the full system end-to-end — ingestion pipeline, embedding strategy, retrieval logic, FastAPI service, and MLflow observability layer. Made all core architecture decisions including chunking strategy, vector DB selection, and LLM orchestration approach.

Key Outcomes

Sub-200ms query latency. Zero in-domain hallucinations — every answer cites its source chunk. First-ever LLM cost dashboard giving visibility into token usage and cost-per-query. Stateless architecture ready for horizontal scaling.

System Architecture — Full Pipeline
📄 Ingestion
PDF / Docs / Wiki→Text Extractor→Semantic Chunker→Embedding Model→Vector DB
🔍 Query
User Query→Query Embedder→FastAPI Service→Top-K Retriever→LLM + Context→Response + Citations
📊 Observe
MLflow Tracker+SQL Metrics Store+Ops Dashboard+Cost / Token Alerts
🔒 Security
Auth Middleware+Rate Limiter+Audit Log+Source Traceability IDs
Ingestion Pipeline Design

Documents are split using semantic chunking to preserve paragraph coherence. Each chunk is embedded and stored with metadata: source file, page number, timestamp, and chunk ID for precise citations at query time.

Retrieval Strategy

Query-time embedding compared via cosine similarity. Top-K=20 retrieved, re-ranked by cross-encoder to top-5. Prompt template enforces the model only answers from provided context — eliminating hallucination.

FastAPI Service Layer

Stateless async FastAPI with UUID traceability per request. Async endpoints prevent blocking on LLM calls. Middleware handles auth, rate limiting, and request timing — feeding every metric into the observability layer.

Four deliberate architecture decisions that shaped the system.

❌ Considered
Fine-tuning an LLM on internal docs

Expensive GPU compute, long cycles, retraining every time docs change. Not viable for a dynamic corpus.

✓ Chose Instead
RAG with Vector DB retrieval

No retraining needed — add a doc, re-embed it. Real-time updates, lower cost, same accuracy.

 

❌ Considered
Fixed-size text chunking (512 tokens)

Fast but splits mid-sentence, breaking semantic coherence. Hurts retrieval quality.

✓ Chose Instead
Semantic chunking by paragraph boundary

94% vs ~78% retrieval accuracy in internal benchmarks.

 

❌ Considered
Synchronous LLM calls in request path

Blocks FastAPI worker threads during LLM latency. Collapses throughput under load.

✓ Chose Instead
Async endpoints + non-blocking LLM client

10× higher concurrent request handling.

 

❌ Considered
Single monolithic service

Ingestion and query have different scaling needs. Can't scale independently.

✓ Chose Instead
Separate ingestion + query services

Query scales independently during peak. Each service has its own failure domain.

01
Retrieval quality degraded past 5K documents

Noisy top-K results — semantically similar but contextually irrelevant chunks slipping into the retrieved set, confusing the LLM.

Added cross-encoder re-ranking on top-20 before trimming to top-5. Accuracy: ~81% → 94%.
02
P99 latency was 3× P50

LLM API tail latency hit 1.8s vs 600ms median — making SLA commitments impossible.

Streaming responses start rendering at first token (~180ms). Parallel async retrieval cut pre-LLM latency to <20ms.
03
Stale embeddings returned outdated policy answers

Updated documents kept old embeddings in Vector DB — users got answers citing superseded versions.

SHA-256 fingerprinting detects changes and triggers selective re-embedding — only modified chunks reprocessed.
04
Token costs invisible until the monthly bill

No real-time cost tracking — no way to know which queries were expensive or if prompt changes were cost-effective.

Every LLM call logged with token counts, cost estimate, latency → SQL + MLflow dashboard. Enabled 30% prompt optimisation.
# FastAPI async endpoint — retrieval + LLM + observability @app.post("/query") async def query_documents(request: QueryRequest, user: User = Depends(auth)): request_id = str(uuid4()) start = time.perf_counter() query_vec = await embedder.aembed(request.query) candidates = vector_db.similarity_search(query_vec, k=20) top_chunks = cross_encoder.rerank(request.query, candidates, top_n=5) prompt = build_prompt(query=request.query, chunks=top_chunks, instruction="Answer ONLY from context. Cite source IDs.") response, tokens = await llm.astream(prompt) latency_ms = (time.perf_counter() - start) * 1000 log_metrics(request_id=request_id, latency_ms=latency_ms, tokens=tokens, cost=estimate_cost(tokens), retrieval_score=top_chunks[0].score) return QueryResponse(answer=response, sources=[c.source_id for c in top_chunks], latency_ms=round(latency_ms, 1))
Why async matters

LLM calls take 500ms–2s. Sync blocks the worker thread. Async + await serves other requests while waiting — 10× throughput improvement under concurrent load.

The grounding instruction

"Answer ONLY from context. Cite source IDs." — forces the LLM to stay within retrieved context. If the answer isn't there, it says so rather than inventing one.

02 Backend · Distributed Systems

Real-Time Log
Analytics Pipeline

Designed and deployed an event-driven Kafka + Airflow distributed pipeline processing millions of log events/day with fault-isolated DAG stages. Optimised PostgreSQL via composite indexing and execution plan tuning — 35% latency reduction, 40% manual ETL ops eliminated.

35%Query latency ↓
40%Manual ops ↓
M+Events / day
KafkaAirflowPostgreSQLPythonETL
GitHub
Pipeline Architecture
Log Sources
→
Kafka Topics
→
Airflow DAGs
→
PostgreSQL
Transform
+
Validate
+
Persist
→
Analytics
Fault-isolated stages · consumer-lag monitoring · dead-letter queue
Problem

Log data from multiple services arrived unstructured and at high volume. Manual processing caused delays, data loss, and 40% ops overhead — no scalable way to query historical events.

Architecture

Event-driven pipeline: Kafka topics per log source → Airflow DAGs for fault-isolated transformation stages → PostgreSQL with composite indexes for analytical queries.

Key Decision

Chose Kafka over direct DB writes for decoupling — producers don't block on consumer lag. Dead-letter queue catches failed events without pipeline stalls. Consumer-lag monitoring via custom Airflow sensors.

Impact

35% query latency reduction via schema redesign and EXPLAIN ANALYZE tuning. 40% manual ops eliminated through automated DAG orchestration. Millions of events/day processed reliably.

03 Machine Learning · Production

Customer Churn
Prediction System

Built end-to-end ML engineering pipeline — Pandas feature engineering → Scikit-learn classification → cross-validation (~87% precision, F1-scored) → Joblib serialisation to deploy-ready .pkl artifact. Sub-50ms inference; wrappable in FastAPI for production serving.

~87%Precision
<50msInference
PklDeploy artifact
PythonScikit-learnPandasNumPyJoblib
GitHub
ML Pipeline
Raw CSV
→
Feature Eng.
→
Classifier
→
.pkl Artifact
Train / Test
→
Eval Metrics
→
Inference API
Retrain pipeline · batch + real-time inference · drift detection ready
Problem

Business needed to identify customers at risk of churning before they left. No existing ML pipeline — raw CSV data, no feature engineering, no model, no deployment path.

Architecture

Pandas ETL for feature engineering → Scikit-learn classifier with cross-validation → Joblib serialisation to .pkl → FastAPI-wrappable inference endpoint. Full retrain pipeline included.

Key Decision

Chose Scikit-learn over deep learning — tabular data, interpretability required, fast inference needed. Decoupled preprocessing from training so retraining only rebuilds affected stages, not full pipeline.

Impact

~87% precision on held-out test data, F1-scored to handle class imbalance. <50ms inference latency. Deploy-ready .pkl artifact wrappable in any Python REST framework in under 10 lines.

04 Data Analytics · Visualisation

Amazon Prime Air
Feasibility Dashboard

Built four-module data pipeline system — Pandas ETL → geospatial feasibility scoring (14 MA counties) → CO₂ emissions modelling → real-time delivery simulator with live Plotly KPI computation. Fully deployable to Streamlit Cloud.

14Counties mapped
4Analysis modules
3×Delivery methods
PythonStreamlitPlotlyPandasGeospatial
GitHub
Dashboard Modules
CSV Datasets
→
Pandas ETL
→
Geo Scoring
→
Choropleth
Emissions Data
→
CO₂ Calc.
→
Plotly Charts
Real-time sliders · live KPI cards · Streamlit Cloud deployable
Problem

Amazon Prime Air feasibility data was scattered across CSV files with no unified analysis layer. Needed a system to score delivery viability, compare emissions, and simulate delivery times across geographies.

Architecture

Pandas ETL pipelines feed four independent modules: geospatial scoring engine, product compatibility filter, CO₂ emissions calculator, and real-time delivery simulator — all rendered via Plotly + Streamlit.

Key Decision

Chose modular Pandas pipelines over a monolithic script — each analysis module is independently composable. Adding a new delivery method or county requires one CSV file swap, no code changes.

Impact

14 Massachusetts counties scored and visualised. 3 delivery methods compared on emissions and time. Real-time KPI computation with zero backend latency. Streamlit Cloud deployable.

Research & AI Projects Published View GitHub ›
🧠
Sahai — AI Mental Health Companion
Gemini AI
🌾
Smart Crop Management — WSN & AI
IEEE Published
🤖
Enterprise GenAI Control Tower
Case Study ↓
Top 10 Projects This Week Trending
1
🤖
GenAI Control Tower
Recently added
2
🌾
Smart Crop Mgmt
IEEE 2023
3
🧠
Sahai AI
Recently added
4
🚁
Prime Air Dashboard
Recently added
5
📉
Churn Prediction
Recently added
6
⚡
Log Analytics Pipeline
Work Project
Engineering Highlights Backend · Data · ML
Distributed pipelines — Designed Kafka + Airflow event-driven systems processing millions of records/day with fault-isolated stages
REST API development — Built scalable FastAPI and Flask services with async endpoints, auth middleware, and rate limiting
Database engineering — PostgreSQL schema redesign, composite indexing, execution plan tuning — 35% latency reduction
ML pipeline deployment — End-to-end training → evaluation → Joblib serialisation → production inference API
Cloud-native ETL — AWS S3, Lambda, Redshift ingestion pipelines eliminating 40% of manual data operations
RAG / LLM systems — Vector DB indexing, cross-encoder re-ranking, async LLM serving with full observability
Performance optimisation — Async FastAPI, query tuning, streaming responses — P99 latency from 1.8s to <200ms
Published research — IEEE ICCCNT 2023 @ IIT Delhi — ML-driven smart agriculture system
Technical Stack Skills
🐍
Python
Expert
☕
Java
Expert
🟨
JavaScript
Advanced
🔷
TypeScript
Advanced
⚡
Kafka
Expert
🌬️
Airflow
Expert
☁️
AWS
Advanced
🗄️
PostgreSQL
Expert
🔴
Redshift
Advanced
🐳
Docker
Advanced
🤖
RAG / LLMs
Advanced
🐳
Docker
Advanced
⚙️
FastAPI / Flask
Expert
🌱
Spring Boot
Intermediate
🧠
PyTorch / TF
Intermediate
Certifications Verified
Global Markets Sales & Trading Analyst Simulation
Bank of America
Cloud Platform Job Simulation — Developer Program
Verizon Communications Inc.
Machine Learning & Big Data Using Python
Professional Certification
Data Science with R Language
Professional Certification
Strategy Consulting Virtual Experience
Virtual Program
About Me
Atmika Manoj Parey

Atmika Manoj Parey

Software & Data Engineer

📍 Boston, MA

she / her

ENGINEER WHO SHIPS AT SCALE

I'm a software engineer with a focus on backend systems, distributed data pipelines, and cloud-native architectures. I've built systems that process millions of records/day, reduced PostgreSQL query latency by 35%, cut manual data ops by 40%, and served REST APIs with 30% faster response times.

Currently finishing my M.S. in Business Analytics at UMass Amherst (graduating May 2025), I bridge the gap between engineering and data-driven decision making. I'm actively exploring opportunities in software engineering, backend, data engineering, and ML engineering.

2024 – MAY 2025
M.S. Business Analytics
University of Massachusetts Amherst · Isenberg School
MAY 2024 – AUG 2024
Software Engineering Intern
DataMind Analytics · Remote
JUN 2022 – DEC 2023
Software Developer
NetCast Services · Mumbai
JUN 2021 – NOV 2021
Software Developer Intern
Supa Ventures Pvt Ltd
2019 – 2023
B.Tech, Computer Science & Business Systems
SRM Institute of Science and Technology
<200msEnd-to-end query latency
10K+Documents indexed
94%Retrieval accuracy
0In-domain hallucinations
3Pipeline stages
The Problem

Enterprise teams spend hours searching unstructured internal documents — PDFs, wikis, reports. Keyword search returns noise with no context. Generic LLMs hallucinate without domain grounding. The gap: fast, accurate, cited answers from private knowledge bases.

The Solution

A two-layer RAG system: an ingestion pipeline that chunks, embeds, and stores documents in a Vector DB, and a query pipeline that retrieves semantically relevant context and passes it to the LLM — ensuring grounded, cited responses with full observability.

My Role

Designed and built the full system end-to-end — ingestion pipeline, embedding strategy, retrieval logic, FastAPI service, and MLflow observability layer. Made all core architecture decisions including chunking strategy, vector DB selection, and LLM orchestration approach.

Key Outcomes

Sub-200ms query latency. Zero in-domain hallucinations — every answer cites its source chunk. First-ever LLM cost dashboard giving visibility into token usage and cost-per-query. Stateless architecture ready for horizontal scaling.

System Architecture — Full Pipeline
📄 Ingestion
PDF / Docs / Wiki → Text Extractor → Semantic Chunker → Embedding Model → Vector DB
🔍 Query
User Query → Query Embedder → FastAPI Service → Top-K Retriever → LLM + Context → Response + Citations
📊 Observe
MLflow Tracker + SQL Metrics Store + Ops Dashboard + Cost / Token Alerts
🔒 Security
Auth Middleware + Rate Limiter + Audit Log + Source Traceability IDs
Ingestion Pipeline Design

Documents are split using semantic chunking (not fixed-size) to preserve paragraph coherence. Each chunk is embedded using a sentence-transformer model and stored in the Vector DB with metadata: source file, page number, timestamp, and chunk ID. This enables precise citations at query time.

Retrieval Strategy

Query-time embedding is computed and compared against stored vectors using cosine similarity. Top-K=5 chunks are retrieved, re-ranked by relevance score, then passed as context to the LLM. The prompt template enforces that the model only answers from the provided context — eliminating hallucination on in-domain questions.

FastAPI Service Layer

Stateless FastAPI handles all query routing. Async endpoints prevent blocking on LLM calls. Each request gets a UUID for full traceability through the audit log. Middleware handles auth, rate limiting, and request timing — feeding every metric into the MLflow + SQL observability layer.

Every major architecture decision involved a deliberate tradeoff. Here are the four choices that most shaped the system.

❌ Considered
Fine-tuning an LLM on internal docs

Would require expensive GPU compute, long training cycles, and retraining every time docs change. Not viable for a dynamic document corpus.

✓ Chose Instead
RAG with Vector DB retrieval

No retraining needed — add a doc, re-embed it. Real-time updates, lower cost, same accuracy on in-domain questions.

 

❌ Considered
Fixed-size text chunking (512 tokens)

Fast and simple, but splits mid-sentence, breaking semantic coherence. Retrieved chunks lack context boundaries — hurts retrieval quality.

✓ Chose Instead
Semantic chunking by paragraph boundary

Preserves natural language units. 94% retrieval accuracy vs ~78% with fixed chunking in internal benchmarks. Worth the extra preprocessing cost.

 

❌ Considered
Synchronous LLM calls in request path

Simplest to implement but blocks FastAPI worker threads during LLM latency (500ms–2s). Under load, this collapses throughput.

✓ Chose Instead
Async endpoints + non-blocking LLM client

10× higher concurrent request handling. FastAPI workers free immediately while awaiting LLM response. No degradation under moderate load.

 

❌ Considered
Single monolithic service

Easier to deploy initially, but ingestion and query workloads have completely different scaling needs — you don't want to scale the whole system when only queries spike.

✓ Chose Instead
Separate ingestion + query services

Query service scales independently during peak usage. Ingestion runs as a background batch job. Each service has its own failure domain.

01
Retrieval quality degraded as the corpus grew past 5K documents

With more documents, the Vector DB returned increasingly noisy top-K results — semantically similar but contextually irrelevant chunks were slipping into the retrieved set, confusing the LLM and producing vague answers.

Added a re-ranking layer using cross-encoder scoring on the top-20 candidates before trimming to top-5. Retrieval accuracy jumped from ~81% to 94%.
02
LLM response latency was non-deterministic — P99 latency was 3× P50

The LLM API had high tail latency — median responses came back in 600ms, but the 99th percentile hit 1.8s. This made SLA commitments impossible and user experience inconsistent.

Implemented streaming responses so the UI starts rendering at first token (~180ms). Parallel async retrieval reduced pre-LLM latency to <20ms. Added a 2s timeout with graceful fallback messaging.
03
Embedding stale documents caused answers to reference outdated policy

Documents were updated frequently, but the Vector DB retained old embeddings. Users received answers citing superseded versions of internal policies — a serious trust problem in an enterprise context.

Built a document fingerprinting system (SHA-256 hash on content) that detects changes and triggers selective re-embedding — only modified chunks are re-processed, keeping ingestion costs low.
04
Token costs were invisible until the monthly API bill arrived

Without real-time cost tracking, there was no way to know which query patterns were expensive, which users were over-consuming, or whether prompt engineering changes were cost-effective.

Instrumented every LLM call with token counts (prompt + completion), cost estimate, and latency — all logged to SQL and surfaced in the MLflow dashboard. Cost-per-query visibility enabled a 30% prompt optimization.
Core RAG Query Handler — FastAPI + LangChain
# FastAPI async endpoint — retrieval + LLM call + observability @app.post("/query") async def query_documents(request: QueryRequest, user: User = Depends(auth)): request_id = str(uuid4()) start = time.perf_counter() # 1. Embed query vector query_vec = await embedder.aembed(request.query) # 2. Retrieve top-K candidates, re-rank with cross-encoder candidates = vector_db.similarity_search(query_vec, k=20) top_chunks = cross_encoder.rerank(request.query, candidates, top_n=5) # 3. Build grounded prompt — LLM only answers from context prompt = build_prompt(query=request.query, chunks=top_chunks, instruction="Answer ONLY from context. Cite source IDs.") # 4. Stream LLM response (non-blocking) response, tokens = await llm.astream(prompt) # 5. Log to MLflow + SQL — latency, tokens, cost, user, request_id latency_ms = (time.perf_counter() - start) * 1000 log_metrics(request_id=request_id, latency_ms=latency_ms, tokens=tokens, cost=estimate_cost(tokens), retrieval_score=top_chunks[0].score) return QueryResponse(answer=response, sources=[c.source_id for c in top_chunks], latency_ms=round(latency_ms, 1), request_id=request_id)
Why async matters here

LLM API calls take 500ms–2s. Synchronous calls block the entire FastAPI worker thread during that wait. Async + await lets the worker serve other requests while waiting, giving 10× throughput improvement under concurrent load with zero extra infrastructure.

The grounding instruction

The key to zero hallucinations: "Answer ONLY from context. Cite source IDs." This single instruction in the prompt template forces the LLM to stay within retrieved context — if the answer isn't there, it says so rather than inventing one.

For Hiring Managers

WHAT I BRING
TO YOUR TEAM

I'm targeting Backend Engineer, Data Engineer, and ML Engineer roles. Here's the quick picture — what I build, how I work, and the impact I've shipped.

35% PostgreSQL query
latency reduced
40% Manual data ops
eliminated
<200ms RAG system
query latency
3+ Years industry
experience
⚙️
What I Build
  • REST APIs — FastAPI, Flask, microservice architecture, async endpoints
  • Streaming pipelines — Kafka topic partitioning, Airflow DAG orchestration, ETL/ELT
  • Database engineering — PostgreSQL schema design, indexing strategies, query optimisation
  • ML engineering — model training, evaluation, Joblib serialisation, production deployment
  • Cloud infrastructure — AWS S3, Lambda, Redshift; Docker; CI/CD pipelines
🎯
How I Work
  • Ship measurable outcomes, not just features
  • Design for observability first — latency, cost, errors tracked from day one
  • Comfortable owning a system end-to-end — design to deployment
  • Agile, code-review culture — clean commits, documented decisions
  • M.S. Computer Science focus — strong foundation in systems, algorithms, and distributed computing
🚀
Key Impact Areas
  • Performance engineering — 35% DB latency cut via schema + index redesign
  • Automation — 40% reduction in manual ops through pipeline automation
  • AI observability — First LLM cost dashboard for token usage + quality
  • Published research — IEEE ICCCNT 2023, IIT Delhi
  • 30% reporting efficiency — Power BI dashboards at DataMind Analytics
Industries I've Worked In
📊 Data & Analytics
☁️ Cloud Infrastructure
🤖 AI / Machine Learning
🌾 AgriTech (IEEE Research)
🏥 Mental Health Tech
📦 Logistics & Delivery
🔧 Backend Systems
📈 Business Intelligence

OPEN TO FULL-TIME ROLES

Looking for full-time opportunities in Software Engineering, Backend, Data Engineering, and ML Engineering.

📋 pareyatmika@gmail.com — copied!