Sparsh Verma
Data Engineer & Cloud Automation Specialist
I design resilient data pipelines and cloud-native workflows on GCP — turning manual, error-prone reporting into systems that run themselves. Over 2+ years shipping production pipelines, BI platforms, and — most recently — applied GenAI automation.
Where the data starts
I'm a data engineer based in Delhi NCR, India, currently building automation and AI-driven data pipelines at HCLTech. My work sits at the intersection of cloud infrastructure, ETL design, and — more recently — applying LLMs to real production problems like PII redaction.
Before this, I cut my teeth at Blackcoffer Consulting, shipping BI dashboards and scraping pipelines for clients across sectors. I like systems that remove a human from a repetitive loop — and I like it even more when I can prove it with a number.
Outside the stack, I led my college's Editorial Club and served as Student Council Coordinator — same instinct for turning chaos into a clean pipeline, just with people instead of data.
- Location
- Delhi NCR, India
- Education
- B.Tech, Computer Science — Galgotias College of Engineering & Technology (2019–2023)
- Focus
- Cloud-native pipelines + applied GenAI
- Status
- Open to new opportunities
Core Expertise
Cloud & IaC
ETL & Processing
BI & Data Ops
AI & GenAI Tooling
Professional Experience
Software Engineer
Jan '24 — PresentHCLTech, Noida, India
- Led a 7-person team delivering a deterministic AI/LLM pipeline that replaced a fully manual content-review step.
- Built automated Python → BigQuery → Power BI reporting pipelines, cutting manual reporting effort by 35%.
- Shipped a PowerApps / Power Automate / FastAPI survey tool to 10,000+ coworkers via SMTP, cutting manual effort by 60%.
- Designed CI/CD workflows for GCP Cloud Run functions, cutting deployment time by 60%.
- Implemented Terraform-based disaster recovery for GCP, enabling one-click infrastructure rebuild and 100% business continuity.
Data Science Intern
Aug '23 — Jan '24Blackcoffer Consulting Firm, Delhi, India
- Built FastAPI pipelines connecting Google Looker Studio to BigQuery, improving reporting speed by 30%.
- Designed and published interactive Power BI / DAX dashboards for client stakeholders.
- Automated Python ETL scripts for data ingestion from external APIs (RapidAPI, Looker API).
- Built automated web-scraping workflows (BeautifulSoup) feeding BigQuery and the ELK Stack for analysis.
Key Engineering Projects
Six production builds — one shipped solo end-to-end, the rest spanning applied GenAI, BI platforms, and cloud-native ETL — each one designed to take a slow, manual process and make it run on its own.
SSBForge — SSB Interview Prep Platform
Live — personal projectA free, full-stack platform helping India's defence aspirants prepare for the SSB interview — six timed practice modules, Supabase-backed auth, and a ₹51/month premium tier, built and shipped solo.
Forge Your Officer's Mindset
Technologies
Key Highlights
- • Six timed practice modules (TAT, WAT, PPDT, SRT, Lecturette, GPE) that mirror real SSB conditions down to the second.
- • Built a Supabase Auth system with session caching and token refresh, gating premium content behind signed URLs.
- • Automated content-index regeneration via GitHub Actions, so the repo can stay public without leaking premium assets.
Deterministic AI/LLM PII Redaction Pipeline
Private — built at HCLTechA hybrid rule-based + LLM pipeline that scrubs PII, profanity, and sensitive content from text at scale — replacing a manual review step for a 7-person team.
Technologies
Key Highlights
- • Combined deterministic rule-matching with LLM judgment to reach 99% scrub accuracy.
- • Removed a fully manual review step, freeing a 7-person team for higher-value work.
- • Built to be deterministic by design — identical input always yields identical redactions, which matters for compliance.
← swipe to see the full pipeline →
Interactive BI Dashboard — Crimes in India
Live demoA live business-intelligence report, fed by an automated GCP + Python pipeline, surfacing actionable insight into NCRB crime data (2020–2022).
Technologies
Key Highlights
- • Connects directly to GCP BigQuery through an automated data pipeline.
- • Built to shorten the path from raw data to a stakeholder decision.
- • Demonstrates end-to-end ELT/BI delivery, from source to visualization.
Client Data Automation Solutions
Private — built at HCLTechAutomated enterprise reporting pipelines, deployed end-to-end on GCP.
Technologies
Key Highlights
- • Streamlined the data flow from SharePoint → BigQuery → Power BI, cutting manual reporting by 35%.
- • Designed CI/CD for Cloud Run functions, cutting deployment time by 60%.
- • Automated the data-freeze process, taking it from 4 hours down to 30 minutes.
← swipe to see the full pipeline →
Automated UK Government Data ETL
Private — built at BlackcofferAn ETL tool that ingests and refines large public datasets for a client-facing web application.
Technologies
Key Highlights
- • Mines large datasets from multiple UK government sources.
- • Refines and structures data for loading into PostgreSQL behind a client-facing app.
- • Keeps stakeholder reporting close to real time.
← swipe to see the full pipeline →
Content Sentiment Analysis Pipeline
Private — built at BlackcofferAn automated script that scores sentiment across 200+ scraped websites.
Technologies
Key Highlights
- • Scrapes content from 200+ sources on an automated schedule.
- • Applies NLP scoring to calculate sentiment per source.
- • Feeds a continuous pipeline for trend analysis and content strategy.
← swipe to see the full pipeline →
Let's ship something
I'm currently a Software Engineer at HCLTech, and always open to conversations about Data Engineering, Cloud Architecture, or applied GenAI roles. Feel free to reach out.