Hey, I'm Aayush 👋

Senior Backend Engineer

Building AI-powered platforms and scalable data systems

Senior Backend Engineer with 5+ years building scalable SaaS platforms, distributed data pipelines, and production ML/LLM systems. I work across Python services, REST and inference APIs, cloud data platforms, GenAI/RAG workflows, MCP, OCR extraction, and resilient enterprise systems.

5+ years Backend engineering GenAI/RAG Distributed systems Microservices System design
Aayush Kumar Mishra

System Architecture

DAG workflows, async queues, caching layers, and service-bus orchestration for resilient production systems.

Backend Engineering

Python microservices, REST and inference APIs, Django, FastAPI, and scalable data-service platforms.

AI/Data Platforms

Bedrock, RAG, OCR extraction, vector databases, Databricks, Snowflake, and ML-driven workflows.

Production Ownership

CI/CD, Kubernetes, observability, release governance, schema safeguards, and reliability-first delivery.

About Me

I build the seam between data platforms and the models that run on them

Backend architecture · Distributed data · Applied AI

I design systems, not just services. Most of what I own starts as a question about boundaries: what belongs in a request path and what belongs behind a queue, where state is allowed to live, which failures a caller should ever see. I work mainly in Python and Scala, across distributed data pipelines, service architecture, and the orchestration that holds the whole thing together.

The stack follows from those decisions rather than the other way round. Kubernetes and Docker for the runtime, AWS and Azure underneath, Kafka, SQS and Service Bus wherever work should be asynchronous, and Snowflake, Databricks, PostgreSQL, MongoDB, Redis or Cosmos DB depending on the access pattern. What I actually spend my judgement on is idempotency, caching strategy, back pressure and observability, because those are what decide whether a system survives its second year.

A growing part of that architecture is LLM inference. I treat a model as one more unreliable dependency with a latency budget and a failure mode: batching and caching to keep inference cost predictable, streaming where someone is waiting, fallbacks for when a provider degrades, guardrails on what a model is allowed to return, and evals that run before a prompt change ships rather than after. I build RAG pipelines, document extraction workflows and agentic systems where the model does the fuzzy reasoning and deterministic services own everything that has to be correct. Knowing exactly where to put that boundary is the real engineering work.

5+
Years Experience
20+
Projects Completed
15+
Technologies

What I Do

Backend Development

Designing scalable APIs, microservices, and robust backend architectures using modern technologies.

AI/ML Engineering

Building intelligent systems with machine learning, deep learning, and generative AI technologies.

Cloud Architecture

Implementing cloud-native solutions on AWS and Azure with focus on performance and scalability.

Currently Building

Referral intelligence for healthcare workflows

Building a configurable healthcare data pipeline that handles fax, email, and manual intake; runs Gemini OCR and Bedrock-backed extraction; writes editable metadata back into Cosmos DB; and drives downstream DAG stages for benefit verification, prior authorization, provider and patient communication, autonomous calling, and scheduling.

Referral intake to scheduling pipeline Fax, email and manual intake feed a Service Bus queue. Configurable extraction runs Gemini OCR and AWS Bedrock, results persist to Cosmos DB as editable metadata, a human review step gates the output, and a DAG fans out to benefit verification, prior authorization, provider communication, autonomous calling and scheduling. Fax Email Manual Queue Service Bus Extract CONFIGURABLE Gemini OCR AWS Bedrock Persist Cosmos DB Review Human in loop Benefit verification Prior authorization Provider comms Autonomous calling Scheduling

Config-driven referral pipeline from intake sources through orchestration, Gemini OCR, Bedrock extraction, Redis enrichment, Cosmos DB writeback, review UI, benefit verification, prior authorization, provider and patient communication, and scheduling.

Experience

Jun 2026 - Present

SDE III

SkyPoint

Taking ownership of engineering responsibilities across healthcare and senior-living applications, with a focus on data modules, platform workflows, and full-stack systems that support data-driven product capabilities, including referral extraction workflows powered by Bedrock-integrated data extraction systems.

  • Streamlined the end-to-end deployment lifecycle, strengthening release governance, CI/CD, environment readiness, and multi-cloud deployment practices.
  • Architected scalable platform workflows using Redis TTL caching, scaled precomputation, service-bus-backed asynchronous processing, and agentic LLM DAG flows.
  • Building referral extraction systems that combine healthcare document workflows, structured data extraction, and AWS Bedrock integration to turn unstructured referral inputs into usable product data.
  • Built data-intensive analytics dashboards with 30+ chart visualizations, interactive drilldowns, optimized queries, and data-quality monitoring for healthcare workflows.
Python React TypeScript REST Cosmos DB Redis Azure ADLS Service Bus Azure DevOps Databricks Kubernetes AWS Bedrock Data Extraction Agentic AI
Jun 2025 - May 2026

SDE - II

Innovaccer Inc.

Built production-grade ingestion, orchestration, and agentic AI workflows for healthcare data platforms, focusing on scalable pipelines, governed data operations, and reliable high-throughput processing.

  • Built a Scala-based dataset pattern recommendation system for ingestion pipelines using Athena metadata, schema-aware file clustering, incremental fetch tracking, partition normalization, and confidence-scored glob generation.
  • Integrated GenAI for pattern inference and conflict resolution with local fallback, existing-pattern context, sample-file validation, retry logic, and manual warnings for risky dataset-in-use conflicts.
  • Implemented production-grade job orchestration with SQS message deduplication, MongoDB job tracking, pod heartbeats, zombie takeover, parallel cluster processing, optimistic locking, atomic updates, and category enrichment.
  • Extended Snowflake-first Scala and Python data pipelines to Databricks with service-principal authentication, PAT fallback, Unity Catalog integration, metastore-aware writes, external locations, governed table operations, and strict date/timestamp safeguards.
  • Designed ingestion enhancements for Kafka metadata size limits through in-payload optimization, compression, chunking, and cloud-backed Avro offloading for oversized metadata.
  • Improved parallel incremental ingestion for S3 and ADLS archive and destination flows, enabling higher-throughput processing and safer handling of large metadata payloads.
  • Integrated multiple agentic workflows across the product lifecycle to support low-latency inference and scalable high-throughput pipelines.
Python Scala Gen AI Snowflake Databricks Athena Kafka PostgreSQL MongoDB SQS Avro AWS Azure S3 ADLS REST Kubernetes ML
Nov 2023 - Jun 2025

Lead Software Engineer : DS/ML

SIRION

Led SaaS development projects across Python REST frameworks, AWS, big data, containerization, document intelligence, and GenAI workflows for enterprise contract platforms.

  • Reduced customer onboarding time by over 85% by delivering service-based integrated SaaS solutions and event-driven webhook designs.
  • Delivered AI/ML-driven document extraction and process automation with generative AI, saving 15,500+ hours and contributing over $2M in additional revenue.
  • Enhanced PDF and scanned-document processing with OCR and ML models, improving processing speed by 30% while persisting extracted fields in PostgreSQL and MongoDB.
  • Built logging, monitoring, and reporting for concurrent processes, reducing index complexity by 40% and improving observability and troubleshooting time.
Python Java LLM Gen AI RAG OCR PostgreSQL MongoDB LangChain AWS Docker Kubernetes
May 2021 - Nov 2023

Software Engineer : Data services

BETSOL

Delivered enterprise data-service products using React, JavaScript, Python, SQL, MongoDB, Bash scripting, and Azure, with a focus on scalable client-facing analytics and reporting systems.

  • Designed and developed product scorecard and data-service applications serving 50+ clients and contributing to over $70M in FY 2022-23 revenue.
  • Reduced API development time by 30% through efficient code management, streamlined reviews, and reusable backend patterns.
  • Built microservices and data-driven insights through direct client engagement, saving 500+ hours across the user base and improving reporting workflows.
  • Optimized API performance and database queries, boosting system efficiency by more than 20%.
  • Reduced security vulnerabilities by 40% while mentoring engineers on code quality, best practices, and optimization.
React Python JavaScript SQL MongoDB MySQL Azure Bash REST Java

Skills & Technologies

AI & Machine Learning

Agentic AI
Self-Reflection Workflows
Machine Learning
Generative AI
MCP
Agentic LLM
RAG
TensorFlow
LangChain
LangGraph
PyTorch

Backend Development

Python
Java
C++
Scala
JavaScript
TypeScript
REST APIs
SOAP
Microservices
GraphQL

Frontend Development

React
TypeScript
Angular.js
three.js
HTML5
CSS3

Data Platforms & Databases

PostgreSQL
Snowflake
Databricks
MongoDB
Cosmos DB
Redis
Pinecone
Neo4j
Vector DBs
MySQL
Oracle DB
InfluxDB
Kafka
Data Pipeline

Cloud & DevOps

AWS
AWS S3
AWS EKS
AWS SQS
AWS Bedrock
Azure
Azure ADLS
Azure Service Bus
Azure DevOps
Azure Container Services
Docker
Kubernetes
Argo CD
CI/CD

Tools, Version Control & Management

Git
GitHub
Jira
Agile
Scrum
Project Management
Kloudfuse
Power Automate
Power BI
BOLD BI
Grafana

Learning (In progress)

Rust
golang
langgraph
MLOps

Certifications

IBM 2024

Machine Learning Specialist - Professional

Issued July 2024

Machine Learning Algorithms
Show credential
Stanford University 2024

Machine Learning Specialization

Issued June 2024

Machine Learning Model Evaluation
Show credential
Stanford University 2024

Advanced Learning Algorithms

Issued June 2024

Neural Networks Advanced ML
Show credential
Stanford University 2024

Unsupervised Learning, Recommenders, Reinforcement Learning

Issued May 2024

Unsupervised Learning Recommenders RL
Show credential
Stanford University 2024

Supervised Machine Learning: Regression and Classification

Issued May 2024

Regression Classification
Show credential
HackerRank 2024

Java Certification

Issued July 2024

Java Data Structures
Show credential
HackerRank 2022

Problem Solving

Issued May 2022

Algorithms Problem Solving
Show credential
IBM 2020

Applied Data Science with Python

Issued September 2020

Data Science Python
Show credential
IBM 2020

IBM Certified Deep Learning Essentials

Issued August 2020

Deep Learning Data Science
Show credential
HackerRank 2020

Python3 Certification

Issued August 2020

Python Programming
Show credential

Let's Connect

I'm always interested in discussing new opportunities, innovative projects, or just having a chat about AI and technology.