Aditya PatelML / DATA ENG

Machine learning and data engineering, from GAN research to cloud-scale deployment.

Open to Opportunities

Machine Learning & AI

Models that survive contact with production.

GANs, diffusion models and LLM systems in PyTorch and TensorFlow, taken past the notebook into deployed inference on AWS SageMaker and GCP Vertex AI.

Generative
WGAN-GP · Conditional DDPM · PaLM-E
Frameworks
PyTorch · TensorFlow · Keras
Serving
SageMaker · Vertex AI · RunPod

Data Engineering

Pipelines built to be boring.

End-to-end ETL orchestrated in Apache Airflow, with automated quality gates and parallel execution across roughly 10 GB of records.

Orchestration
Apache Airflow · BullMQ
Stores
PostgreSQL · DynamoDB · S3 · DuckDB
Scale
~10 GB EHR · 3× throughput

Full-Stack Development

The surface that makes it usable.

Typed Next.js frontends over REST APIs in ASP.NET MVC and Node, with role-based access control where the data needs it.

Frontend
Next.js · TypeScript · React
Backend
ASP.NET MVC · C# · Fastify
Contracts
REST · role-based access control

IoT & Embedded

Where the data actually comes from.

Arduino and Raspberry Pi sensor firmware, local aggregation and real-time dashboards for the hardware that produces the data in the first place.

Hardware
Arduino · Raspberry Pi
Telemetry
Real-time aggregation · anomaly alerts
Shipped
SHEMS energy monitor

Aditya PatelPractice

From researchto production.

Most of the hard problems live outside the model file: data you cannot trust, a pipeline that has to run at 4am, and a deployment nobody wants to be paged about.

4

Industry internships

AI · Data Eng · Web · IoT

9

Projects shipped

ML, ETL, IoT & web

50K+

Synthetic records

SynMedix on SageMaker

84

Instruments modelled

Global Market Lab

Aditya PatelCapabilities

Technical stack.

Everything listed here has shipped in something real: an internship deliverable, a deployed model, or a project running today.

01

Languages

Primary
Python
Backend
C# · C / C++
Web
TypeScript · JavaScript
Query
SQL
02

ML / AI

Frameworks
PyTorch · TensorFlow · Keras
Generative
GANs · Diffusion · LLMs
Classical
Scikit-learn
Multimodal
PaLM-E · cross-attention
03

Data & Cloud

Orchestration
Apache Airflow
Relational
PostgreSQL · SQL Server
NoSQL
DynamoDB · Redis
Object store
AWS S3
04

Web & Systems

Frontend
Next.js · React
Backend
ASP.NET MVC · Fastify · Node
Interfaces
REST APIs
Embedded
Arduino · Raspberry Pi
05

Tooling

Containers
Docker
CI / CD
GitHub Actions · Git
Analysis
Pandas · NumPy
Reporting
Power BI
06

Deployments

AWS
SageMaker · S3 · DynamoDB
GCP
Vertex AI
GPU
RunPod RTX 4090
Edge
Vercel

Featured BuildSynMedix AI

A distributed EHR processing platform and the generative model on top of it. Three steps from raw records to a dataset a research team can train on.

  1. 01The problem

    Ten gigabytes of records nobody could touch

    Large volume, absolute privacy constraints, and serial processing slow enough that iterating on a model means waiting overnight.

    Input
    ~10 GB electronic health records
    Constraint
    No real patient data downstream
  2. 02The pipeline

    Parallel execution, then a generative layer on top

    Ingest restructured around parallel execution, cutting turnaround to a third. A generative layer on the cleaned corpus learns the joint distribution rather than copying any single record.

    Throughput
    3× over the serial baseline
    Stack
    Python · SQL · PyTorch · TensorFlow
  3. 03The outcome

    50,000+ synthetic records, zero real patients exposed

    Deployed on SageMaker, emitting records that keep the statistical structure of the source corpus without carrying any individual through it.

    Generated
    50,000+ synthetic patient records
    Deployment
    AWS SageMaker

Aditya PatelContact

Let's buildsomething.

A hard ML problem, a pipeline that won't scale, or a role where you need someone who will actually dig in. Leave an address and I'll reply within 24 hours.

Prefer the long form? Full contact page