AI-Powered ML Training Analysis Platform
LLM AgentsScroll ↓Overview
A conversational assistant that reads ML training logs, diagnoses problems, and recommends fixes — built on Google Gemini and the Model Context Protocol.
Details
- Role
- Design & development
- Category
- LLM Agents
- Links
- GitHub ↗
- 18MCP diagnostic tools
- 6validated REST endpoints
- SSEstreaming responses
01 — Agentic Architecture
- Designed a React → FastAPI → Gemini + MCP architecture in which Gemini reasons over structured results from a dedicated MCP server exposing 18 diagnostic tools — e.g. overfitting detection, run comparison and scheduler recommendation.
- Implemented multi-turn, context-aware chat with automatic tool invocation and streaming responses over Server-Sent Events (SSE).
02 — Diagnostics & Recommendations
- Parses CSV training logs to find the best epoch, best validation score and convergence point, and detects overfitting, underfitting, plateaus, stagnation and validation instability.
- Recommends learning rate, optimizer, scheduler, batch size and dropout changes, and compares multiple runs side by side to pick the best configuration.
03 — Tooling & Reporting
- Added a dataset analyzer (duplicates, corrupted images, missing labels, class imbalance), GPU telemetry (VRAM, batch-capacity estimate), auto-generated loss / AUC / LR charts and downloadable Markdown reports, served through 6 validated REST endpoints.
Built with
Python / FastAPI / Uvicorn / Pydantic / Pandas / NumPy / Matplotlib / Gemini 2.5 Flash / Pro / MCP Python SDK / React / TypeScript / Vite / Tailwind CSS / shadcn/ui / Framer Motion