Available for Data Engineering Roles

Hi, I'm Satyajeet Dharmadhikari

ETL Developer - Data Engineer • VOIS

Passionate Data Engineer with 3+ years of hands-on experience designing robust ETL pipelines, high-volume data staging, and server-to-server legacy codebase migrations. Google Cloud Certified Professional Data Engineer and completed M.Tech in Cloud Computing at BITS Pilani.

Satyajeet Dharmadhikari
GCP Certified GCP Certified PDE
⚡ VOIS Data Engineer
3+ Years of Industry Experience
GCP Professional Data Engineer
2+ Major Enterprise Migrations
99.9% Pipeline Reliability & SLAs

Bridging Data Infrastructure & Cloud Intelligence

Transforming complex big data challenges into high-efficiency, reliable, and scalable automated pipelines.

I am an ETL Developer - Data Engineer based in Pune, India, currently driving enterprise data solutions as a Senior Executive at VOIS.

My expertise spans the entire data lifecycle: from high-throughput batch ingestion using Ab Initio and Teradata staging environments, to complex SQL query optimization, Unix/Linux server automation, and cloud modernization on Google Cloud Platform and AWS.

I recently spearheaded a crucial server-to-server legacy codebase migration, modernizing data pipelines with zero operational downtime. Completed my M.Tech in Cloud Computing at BITS Pilani in January 2026 while continually expanding my cloud-native lakehouse and distributed engineering capabilities.

📍
Location: Katraj, Pune, MH, India
💼
Company: VOIS (Vodafone Intelligent Solutions)
🎓
Academics: BITS Pilani (M.Tech Cloud Computing, Jan 2026)
ETL & Cloud Architecture Core
ETL Data Pipeline Architecture
Official Google Cloud Certification

Professional Data Engineer

Issued by Google Cloud • Verified on Credly

Demonstrates proven proficiency to design, build, operationalize, secure, and monitor data processing systems on Google Cloud Platform. Certified in architecting robust streaming & batch systems, big data analytics, and machine learning model integration.

Google BigQuery Cloud Dataflow (Beam) Cloud Dataproc (Spark) Cloud Storage Cloud Pub/Sub Cloud Composer (Airflow) IAM & Cloud Security Data Governance

Work Experience

Delivering resilient data systems and enterprise infrastructure for global telecom scale.

Senior Executive — ETL & Data Engineering

VOIS (Vodafone Intelligent Solutions)

2023 — Present
  • Core ETL Development: Architect and maintain production-grade data pipelines processing enterprise-scale telecommunications datasets using Ab Initio and Teradata.
  • Server Migration Leadership: Successfully executed server-to-server legacy codebase migration, validating multi-terabyte data staging environments with zero data loss and minimal pipeline downtime.
  • Performance Tuning & SQL Optimization: Optimized complex Teradata SQL queries, indexing, and data staging schemas, resulting in reduced batch execution windows.
  • Pipeline Automation: Engineered custom Python and Unix Shell automation scripts for automated job dependency checks, data reconciliation, and alerting.
Ab Initio Teradata UNIX / Linux Advanced SQL Python Shell Scripting Migration

Graduate Engineer Trainee

VOIS (Vodafone Intelligent Solutions)

2022 — 2023
  • ETL Unit Testing & Validation: Designed comprehensive ETL test suites and verified data transformations across staging and warehouse layers.
  • Perl & Shell Scripting: Built automated test-validation scripts in Perl and Bash to automate file format verifications, delimiter checks, and log monitoring.
  • Production Support & Issue Resolution: Monitored daily/weekly batch schedules on UNIX servers, analyzed failure logs, and delivered swift defect resolutions.
ETL Testing Perl Bash SQL Teradata Linux

Skills & Architecture Matrix

Technologies and tools I leverage to build scalable, resilient data pipelines.

ETL & Data Pipelines

Designing, scheduling, and orchestrating massive batch & stream data flows.

  • Ab Initio (GDE, Co>Operating System) Advanced
  • Data Staging & Ingestion Expert
  • Data Warehousing (EDW Architecture) Advanced
  • Server-to-Server Codebase Migration Specialist

Cloud & Big Data

Cloud-native data architecture, managed services, and distributed storage.

  • Google Cloud Platform (BigQuery, Storage) Certified
  • Cloud Dataflow & Cloud Dataproc Proficient
  • AWS Cloud Computing (BITS Pilani) Proficient
  • Docker & Containerization Basics Intermediate

Databases & SQL

Writing performant queries, schema design, and analytical warehousing.

  • Advanced SQL Query Optimization Expert
  • Teradata Database & Utilities (BTEQ, FastLoad) Advanced
  • Google BigQuery (Serverless Analytics) Advanced
  • PostgreSQL & MySQL Proficient

Languages & Scripting

Automating workflows, data parsing, testing frameworks, and tooling.

  • Python (Data Analytics, Scripting, Automation) Advanced
  • Linux / UNIX Shell Scripting (Bash) Advanced
  • Perl Scripting & Text Processing Proficient
  • Core Java & C++ Foundation

Key Projects & Systems

Open-source streaming lakehouses, data lineage architectures, and mission-critical enterprise migrations.

Streaming Lakehouse

Financial Market Lakehouse & Streaming Platform

Real-time financial market data platform architected for ultra-low latency analytics. Integrates Apache Beam streaming pipelines, PySpark distributed compute, Apache Iceberg table format via PyIceberg, Arrow Flight high-throughput transport, and embedded DuckDB analytics.

✓ Real-time streaming ingestion with Apache Beam
✓ ACID table lakehouse via PyIceberg
✓ Ultra-fast in-memory queries with DuckDB & Arrow Flight
Apache Beam PySpark PyIceberg Arrow Flight DuckDB Python
Data Observability

Data Lineage & Pipeline Tracking with PySpark

End-to-end data lineage, metadata collection, and observability system for PySpark pipelines. Implements OpenLineage standards to trace schema drift, dataset dependencies, and execution runtimes for enhanced pipeline governance.

✓ OpenLineage standard integration for Spark
✓ Automated schema drift and dependency graphs
✓ Complete auditability for enterprise data flows
PySpark OpenLineage Metadata Governance Python
Time-Series Analytics

Gold Price Prediction using Holt-Winters Exponential Smoothing

Comprehensive statistical time-series modeling project applying Holt-Winters Exponential Smoothing (HWES) on multi-year historical commodity data. Includes trend decomposition, seasonality modeling, hyperparameter optimization, and MAPE evaluation.

✓ Triple Exponential Smoothing (HWES) algorithm
✓ Multi-year seasonality & trend decomposition
✓ Statistical residual error diagnostics
Time Series Statsmodels Pandas Jupyter
NLP & Web Application

Amazon Reviews Sentiment Analysis Web Application

Natural Language Processing web service that processes unstructured e-commerce product reviews to classify customer sentiments into positive, neutral, and negative categories with text tokenization and TF-IDF feature extraction.

✓ NLP text preprocessing & stopword filtering
✓ TF-IDF vectorization & classification model
✓ Interactive web dashboard for live predictions
NLP Python Scikit-Learn Web App
Enterprise Infrastructure
★ VOIS Production

Server-to-Server Legacy ETL Codebase Migration

Spearheaded the migration of legacy Ab Initio ETL graphs, Unix wrapper scripts, and staging data across enterprise server clusters at VOIS. Designed automated checksum reconciliation to ensure 100% data fidelity with zero production downtime.

✓ Zero downtime transition for critical pipelines
✓ Automated multi-TB staging reconciliation
✓ Optimized cron schedule triggers
Ab Initio Teradata UNIX Shell Python
Data Quality & QA
★ VOIS Automation

Automated Data Validation & Staging Reconciliation

Engineered a custom automated data reconciliation tool to validate millions of staging records against target warehouse tables. Replaced manual unit-testing routines with Perl and Python validation scripts, cutting staging cycle times by over 60%.

✓ 60% faster QA verification cycle
✓ Automated anomaly detection & alert logging
✓ Delimiter & schema validation scripts
Python Perl SQL Linux

Education & Credentials

Solid foundations in computer science, software architecture, and modern cloud technologies.

Jan 2024 — Jan 2026 (Completed)

M.Tech in Cloud Computing

BITS Pilani (WILP)

Specialized program focused on Cloud Architecture, Distributed Computing, Virtualization, Big Data Systems, Cloud Security, and Enterprise Scale Infrastructure.

Graduated

B.Tech in Computer Science & Engineering

Bachelor's Degree

Comprehensive training in Data Structures & Algorithms, Database Management Systems (DBMS), Operating Systems, Object-Oriented Software Engineering, and Computer Networks.

Ready to build something impactful?

Whether you have an exciting data engineering role, a cloud lakehouse project, or want to discuss tech — let's connect!

Email Address
Open in Gmail
Location
Katraj, Pune, Maharashtra, India