Charitha

Lehi, Utah

Charitha Sree Veluru

Data Engineer

3+ years of experience building and optimizing end-to-end data pipelines, dimensional models, and warehouse infrastructure in healthcare and enterprise environments. Blending data engineering fundamentals with modern AI — GenAI, RAG pipelines, and agentic systems — to build intelligent, data-driven platforms.

Experience

Agentic AI Engineer

Dec 2025 – Present

Humanitarians AI · Boston, MA (Volunteer)

  • Architected an AI education platform with Python/Flask, LangChain, LangGraph, and LangSmith — multi-agent RAG pipelines over PostgreSQL/PGVector for semantic search across 100K+ research documents at 99.8% content accuracy.
  • Built ingestion pipelines transforming 100K+ unstructured research papers into structured course content, reducing course creation time by 95% while serving 25,000+ students.
  • Developed a Next.js analytics dashboard with LangSmith agent tracing, achieving a 4.9/5 NPS score.
  • Automated AWS ECS deployments via Docker, Terraform, and GitHub Actions CI/CD — 98% uptime, 40% lower infrastructure costs.

Data Engineer Intern

Jan 2025 – May 2025

ClinicMind · Boston, MA

  • Designed ETL pipelines ingesting billing and insurance data from 10+ providers into Amazon Redshift with Apache Airflow and Python, cutting report turnaround from 3 days to same-day.
  • Optimized SQL workloads on 500K+ claims records, identifying revenue leakage and reducing claim rejection rates by 22%.
  • Engineered a GPU-accelerated claim denial prediction model (XGBoost, RAPIDS cuML, SHAP) with 91% accuracy, reducing denial rates by 18%.
  • Deployed a real-time anomaly detection pipeline with TensorFlow Autoencoders, cutting revenue cycle errors by 25% and saving ~$120K annually.

Data Engineer

Jul 2021 – Jul 2023

Wipro Limited · India

  • Built data pipelines for a medical equipment client, ingesting procurement, inventory, and supply chain data into PostgreSQL — eliminating manual workflows behind 40% of order delays.
  • Engineered a multi-warehouse inventory system with Redis and RabbitMQ for real-time sync, driving oversell incidents to 0% and cutting reconciliation from 3 hours to 10 minutes.
  • Replaced 4–6s PostgreSQL LIKE queries with Elasticsearch-backed sub-200ms responses across a 10,000+ SKU catalog (30x faster).
  • Automated CI/CD with GitHub Actions, Docker, and AWS ECS, cutting 4-hour release cycles to under 18 minutes with 100% success rate.

Data Analyst Intern

Jul 2019 – Aug 2019

Bharat Electronics Limited · India

  • Analyzed operational and production datasets with Python and SQL, reducing manual reporting effort by 40%.
  • Built a full-stack KPI dashboard (React.js, Node.js, PostgreSQL), cutting report turnaround from 2 days to 4 hours.

Skills

Data Engineering & ETL

  • Apache Airflow
  • dbt
  • ETL/ELT Pipeline Design
  • Data Warehousing
  • Kimball Star Schema
  • Incremental Loading
  • Data Quality Testing

Databases & SQL

  • PostgreSQL
  • Amazon Redshift
  • Snowflake
  • MongoDB
  • Redis
  • DuckDB
  • Query Optimization
  • Window Functions & CTEs

Programming

  • Python
  • Pandas & NumPy
  • SQL
  • Shell Scripting
  • Java
  • JavaScript/TypeScript
  • Flask

Cloud & Infrastructure

  • AWS (S3, RDS, Redshift, ECS, Lambda)
  • GCP (BigQuery, Dataflow, Pub/Sub)
  • Terraform
  • Docker
  • Kubernetes
  • GitHub Actions CI/CD

AI/ML & GenAI

  • LangChain
  • LangGraph
  • RAG Pipelines
  • PGVector
  • TensorFlow
  • XGBoost
  • Agentic AI
  • NL-to-SQL

Governance & Compliance

  • HIPAA-compliant Architectures
  • PHI Segregation
  • RBAC
  • Audit Logging
  • Data De-identification

Projects

MedMarket Global

Healthcare Data Platform

MedMarket Global on GitHub
  • End-to-end healthcare platform ingesting data from 7 source systems (EHR, CRM, HR, credentialing, call center) via containerized connectors.
  • Kimball-style star schema with 5 dimensions, 3 fact tables, and 28,000+ encounters, including 25 automated dbt data quality tests and SCD Type 2 tracking.
  • 574x query performance improvement (43ms to 0.076ms) via composite indexes and materialized views.
  • HIPAA-style compliance: PHI/reporting schema split, database-enforced RBAC, audit logging, and salted de-identification.
  • AI agent (Claude API) for natural-language-to-SQL queries with dual-layer security.
  • Python
  • PostgreSQL
  • dbt
  • Docker
  • Terraform
  • AWS
  • Anthropic API
  • Streamlit

AbroadIQ

AI Study Abroad Recommendation Platform

AbroadIQ on GitHub
  • Data scraping pipelines ingesting exchange programs, visa requirements, cost-of-living data, student blogs, and Reddit threads into a structured knowledge base.
  • Dual-agent RAG architecture with PGVector semantic search, BM25 hybrid retrieval, and cross-encoder reranking — sub-2s response latency.
  • High-throughput embedding pipeline (Sentence Transformers) enabling real-time semantic search across 100+ sources in under 500ms.
  • Personalized study abroad recommendations with budget estimates and quality-scored advice.
  • Python
  • LangGraph
  • Groq API
  • PGVector
  • Django
  • Docker
  • BeautifulSoup

Education & Certifications

M.S. in Information Systems

Northeastern University · Boston, MA

Dec 2025

B.Tech in Computer Science and Engineering

PES University · Bangalore, India

May 2021

AWS Certified Solutions Architect – Associate

Amazon Web Services · Nov 2025