Hamza Mhedhbi · Ariana, Tunisia

AI Engineer

I build and ship production AI systems end-to-end.

LLMs · RAG · VLMs · NLP · Document Intelligence · AI Infrastructure

What I do

I design and ship production AI systems that turn complex enterprise data into intelligent products — from ingestion and modeling to APIs, UI, and deployment.

Hours → seconds
report generation
8
systems unified into one CRM
98–100%
data coverage, zero downtime

Focus areas

  • LLM applications, RAG & agentic workflows
  • Document intelligence with Vision-Language Models (Qwen-VL, SmolDocling)
  • Real-time ETL & CRM integration at enterprise scale
  • Production deployment on AWS SageMaker, Azure ML & self-hosted VPS

Core stack

PythonFastAPIReactPyTorchHugging FaceLLaMAGPT-4Qwen-VLFAISSRAGLoRAWhisperPostgreSQLDockerAWS SageMakerAzure ML

Professional experience

Five production AI systems I built end-to-end — what they do, what I personally shipped, and the measurable result.

01

AI Factory

Complaint Analytics Platform · Poulina Group Holding

Automated hotel complaint analysis and reporting: complaints are classified by an LLM, action plans are generated automatically, and teams resolve them on a Kanban board.

Hours → seconds
report generation
2 LLMs
classification + action plans
Production
Docker + Nginx on VPS

My contribution

  • Built the LLM classification and action-plan pipeline
  • Designed the FastAPI backend and PostgreSQL data model
  • Built the React Kanban workflow with RBAC + JWT auth
  • Shipped the production deployment (Docker Compose + Nginx)

Tech

LLaMA 3.3 70BGroqFastAPIReactPostgreSQLDocker
Technical details
Multi-Hotel Complaint Analytics Platform
01 — Problem
Multi-hotel complaints were triaged manually — inconsistent categorization, slow response, and hours to produce action-plan reports for management.
02 — Solution
A platform that ingests complaints, classifies them with LLaMA 3.1 8B (batching + retry logic), auto-generates action plans with LLaMA 3.3 70B, and routes tasks through a Kanban board across lifecycle states.
03 — Architecture
FastAPI backend, PostgreSQL for state, React frontend with a Kanban task board, RBAC + JWT authentication for permission-based access, and containerized deployment on an Oxahost VPS using Docker Compose behind Nginx.
04 — Results
  • Report generation reduced from hours to seconds
  • Kanban board links complaints to tasks across their full lifecycle
  • Deployed to production with Docker Compose + Nginx on Oxahost VPS
05 — Full stack
FastAPIReactPostgreSQLGroqLLaMA 3.1 8BLLaMA 3.3 70BJWTRBACKanbanDocker ComposeNginxOxahost VPS
02

Data Integration Tower

CleverTap Integration · Poulina Group Holding

An hourly pipeline that pulls customer data out of eight separate hotel and park systems and keeps one clean, deduplicated CRM view in CleverTap.

8
systems unified
98–100%
record coverage
24/7
zero downtime

My contribution

  • Built the hourly ETL services for 8 source systems
  • Designed the MD5 hash-ledger deduplication engine
  • Wrote the sanitization layer (dates, nulls, E.164 phones)
  • Dockerized and operated the pipeline on an Ubuntu VPS

Tech

PythonFlaskPostgreSQLCleverTapDocker
Technical details
CleverTap CRM Integration
01 — Problem
Customer data lived in 8 disparate systems (2 PMS, events, VR venue, pool, timeshare, park, WiFi) across 5 hotels and 2 parks — synced manually via CSV, blocking unified CRM engagement.
02 — Solution
A fully automated hourly ingestion pipeline that normalizes, deduplicates, and pushes customer events into CleverTap as a single source of truth.
03 — Architecture
Flask ETL services, PostgreSQL staging, an MD5 hash-ledger deduplication engine decoupled from unreliable DB timestamps, and ETL logic handling dates, null/NaN sanitization, and E.164 phone normalization — Dockerized on an Ubuntu VPS.
04 — Results
  • Unified 8 disparate systems across 5 hotels and 2 parks into one CleverTap view
  • Replaced manual CSV exports with hourly real-time sync
  • 98–100% record coverage · 24/7 with zero downtime
05 — Full stack
CleverTapETLFlaskDockerUbuntu VPSPostgreSQLMD5 DeduplicationE.164
03

Document Intelligence Laboratory

OCR Pipeline Upgrade · BinitNS – Accelex

A document-understanding pipeline that reads complex financial documents with Vision-Language Models instead of legacy OCR.

100%
Ollama uptime in production
Real-time
document parsing

My contribution

  • Optimized SmolDocling with patching, overlapping and MSER
  • Deployed Qwen-VL 2.5 on AWS SageMaker via Ollama
  • Wrote the Bash sustainer that keeps inference always-on

Tech

Qwen-VL 2.5SmolDoclingAWS SageMakerOllama
Technical details
VLM-based OCR Pipeline
01 — Problem
Legacy Tesseract-based OCR at Accelex struggled with complex documents, requiring heavy manual review.
02 — Solution
Proposed and prototyped a full OCR pipeline upgrade replacing Tesseract with VLM-powered document intelligence.
03 — Architecture
Optimized SmolDocling performance using patching, overlapping, and MSER; deployed Qwen-VL 2.5 on AWS SageMaker via Ollama; engineered a custom Bash script to maintain Ollama's operational sustainability in production.
04 — Results
  • Real-time parsing with Qwen-VL 2.5 on SageMaker
  • 100% Ollama uptime in production via custom Bash sustainer
  • Proposed VLM upgrade path replacing Accelex's legacy Tesseract stack
05 — Full stack
Qwen-VL 2.5SmolDoclingAWS SageMakerOllamaMSERBounding Box DetectionBash
04

Language AI Campus

LegalBot · French HR Chatbot

Two LLM assistants for real users: a retrieval-based legal Q&A tool and a French-language HR chatbot grounded in live company data.

90%
less Q&A generation time
2 assistants
legal + French HR

My contribution

  • Built the Azure preprocessing + OCR pipeline for legal documents
  • Generated Q&A pairs with GPT-4 and LLaMA 3 8B (Unsloth)
  • Fine-tuned a French HR model with LoRA and added SpaCy NER
  • Implemented FAISS retrieval and the SharePoint data sync

Tech

LLaMA 3 8BLoRAFAISSRAGGPT-4Azure ML
Technical details
LegalBot · Approach AI (10/2024 – 02/2025)
01 — Problem
Legal teams spent hours preprocessing documents and manually building Q&A pairs for downstream retrieval.
02 — Solution
An automated preprocessing + OCR pipeline on Azure Blob Storage, GPT-4 and LLaMA 3 8B (Unsloth) for Q&A generation, and FAISS-backed RAG for retrieval.
03 — Architecture
Legal document preprocessing and OCR on Azure Blob Storage, fine-tuning and orchestration in Azure ML Studio, GPT-4 + LLaMA 3 8B via Unsloth for Q&A pair generation, FAISS index for retrieval.
04 — Results
  • 90% reduction in Q&A pair generation time
  • Reusable Azure ML Studio pipeline for legal corpora
05 — Full stack
Azure Blob StorageAzure ML StudioOpenAI GPT-4LLaMA 3 8BUnslothFAISSRAGOCR
French HR Chatbot · SACEM Industries (06/2024 – 08/2024)
01 — Problem
HR at SACEM Industries received the same French-language policy questions daily with no self-service tool grounded in live SharePoint data.
02 — Solution
A LoRA-fine-tuned French chatbot with NER, RAG over HR policies, and Power Automate sync between SharePoint and Excel Online for real-time data.
03 — Architecture
LLaMA 3 8B fine-tuned with LoRA, SpaCy for French NER, FAISS-backed RAG, SharePoint↔Excel Online sync via Power Automate, and a Streamlit real-time interface.
04 — Results
  • Real-time French-language HR self-service
  • Live data via SharePoint / Excel Online sync
05 — Full stack
LLaMA 3 8BLoRASpaCyNERRAGFAISSSharePointExcel OnlinePower AutomateStreamlit
05

Voice AI Studio

Speaker ID & Transcription · Omnilink

A call-audio pipeline that separates speakers, transcribes what they said, and labels each speaker as employee or customer.

3-stage
diarization → ASR → classification
Market-ready
adopted by Omnilink

My contribution

  • Built audio preprocessing (noise reduction, normalization, segmentation)
  • Integrated pyannote diarization and Whisper transcription
  • Added LLaMA 2 speaker classification and a Flask service layer

Tech

pyannoteWhisperLLaMA 2PyTorchFlask
Technical details
Speaker ID + Transcription
01 — Problem
Call recordings arrived as a wall of unattributed text — no way to know who spoke or whether they were an employee or a customer.
02 — Solution
A pipeline that preprocesses audio, diarizes and transcribes it, and classifies each speaker as employee vs. customer using LLaMA2.
03 — Architecture
Noise reduction, normalization, and segmentation for audio prep; pyannote/speaker-diarization for segmentation; Whisper AI for transcription; LLaMA2 for employee-vs-customer classification; Flask interface for portability and scalability.
04 — Results
  • Attributed transcripts with employee-vs-customer labels
  • Market-ready system Omnilink plans to bring to production
05 — Full stack
pyannoteWhisperLLaMA 2PyTorchFlaskNoise ReductionNormalizationSegmentation

Career timeline

A chronological view of the roles and systems behind the work above.

  1. 10/2025 —
    Poulina Group Holding

    AI Engineer

    AI Factory: multi-hotel complaint analytics with LLaMA 3.1 8B / 3.3 70B, RBAC + JWT, Kanban workflow, and a CleverTap CRM integration unifying 8 systems across 5 hotels and 2 parks.

    • Report generation cut from hours to seconds
    • 8 systems unified into CleverTap — 98–100% coverage, 24/7 zero downtime
    • Deployed with Docker Compose + Nginx on Oxahost VPS
    FastAPIReactPostgreSQLGroqLLaMACleverTapDockerNginxJWT
  2. 02/2025 — 06/2025
    BinitNS – Accelex

    AI Engineer Intern · OCR Pipeline Upgrade

    Replaced legacy Tesseract-based OCR with Vision-Language Models — Qwen-VL 2.5 deployed on AWS SageMaker via Ollama.

    • Real-time parsing.
    • 100% Ollama uptime via custom Bash sustainer
    • Optimized SmolDocling with patching, overlapping, and MSER
    Qwen-VL 2.5SmolDoclingAWS SageMakerOllamaMSERBash
  3. 10/2024 — 02/2025
    Approach AI

    AI Engineer Intern · LegalBot

    Automated legal document preprocessing + OCR on Azure, GPT-4 / LLaMA 3 8B (Unsloth) for Q&A generation, and FAISS-backed RAG.

    • 90% reduction in Q&A pair generation time
    • Reusable Azure ML Studio pipeline for legal corpora
    Azure Blob StorageAzure ML StudioGPT-4LLaMA 3 8BUnslothFAISSRAG
  4. 06/2024 — 08/2024
    SACEM Industries

    AI Engineer Intern · French HR Chatbot

    French-language HR chatbot with LoRA-fine-tuned LLaMA 3 8B, SpaCy NER, FAISS RAG, and SharePoint / Excel Online sync via Power Automate.

    • Real-time French HR self-service on live SharePoint data
    • Streamlit interface deployed for daily HR use
    LLaMA 3 8BLoRASpaCyFAISSSharePointPower AutomateStreamlit
  5. 06/2024 — 08/2024
    Omnilink

    AI Engineer Intern · Speaker ID & Transcription

    Audio pipeline combining pyannote diarization, Whisper transcription, and LLaMA2 employee-vs-customer classification, packaged behind Flask.

    • Attributed transcripts with employee-vs-customer labels
    • Market-ready system Omnilink plans to bring to production
    pyannoteWhisperLLaMA 2PyTorchFlask
  6. 07/2023 — 08/2023
    Tunisie Telecom

    ML Intern · SMS Spam Detection

    NLP spam classifier benchmarking Multinomial Naive Bayes, XGBoost, and fine-tuned BERT over TF-IDF, CountVectorizer, and GloVe features.

    • Comparative benchmark of classical ML and transformer approaches
    Pythonscikit-learnNaive BayesXGBoostBERTTF-IDFGloVe

Projects

Every project has its own tower in the AI city — or open any of them directly from the list below.

Experiments & lab

Smaller side projects I build to stay sharp — implementations from scratch, learning tools, and open-source experiments.

NeetCode Visualizer

Animated walkthroughs for all 150 NeetCode problems, built with Django.

View on GitHub ↗

Transformer Lab

A GPT-style character-level transformer language model from scratch.

View on GitHub ↗

AI/ML Flashcards Station

1,721 flashcards from three AI/ML books, served via GitHub Pages.

View on GitHub ↗

Machine Learning Museum

Classical ML fundamentals — Naive Bayes, XGBoost, fine-tuned BERT.