Turning Messy Datasets Into
Executive Intelligence
DataForge is an automated data science platform engineered to eliminate manual spreadsheet bottlenecks. From dataset cleaning and statistical insight discovery to AutoML predictive modeling and natural language AI queries—DataForge transforms raw files into decision-ready executive reports in seconds.
Why We Engineered DataForge
Every single day, organizations, researchers, and students generate massive volumes of spreadsheet data. Yet, turning raw rows into meaningful answers remains a tedious, error-prone chore. Traditional software forces users into manual formula crafting, while advanced data science tools demand expertise in Python, R, or SQL.
We built DataForge to eliminate this technical divide. Conceived as both an academic breakthrough and a practical engineering solution, DataForge unifies automated data preprocessing, statistical insight discovery, predictive machine learning, and AI chat into a seamless web platform.
By automating the heavy lifting—from missing value imputation and Z-score outlier detection to FLAML model training and WeasyPrint executive PDF generation—DataForge empowers non-technical users and data analysts alike to derive enterprise-grade insights in seconds.
The Challenges We Solve
Traditional data workflows break down due to manual bottlenecks, fragmented software, and complex statistical code.
Corrupted & Messy Files
Raw CSV and Excel files contain missing values, duplicate entries, inconsistent date strings, and bad data types that corrupt analytics.
Manual & Slow Workflows
Analysts spend over 80% of their time writing repetitive formulas, filtering columns manually, and copying charts into slide decks.
Steep Technical Barriers
Advanced statistical modeling and machine learning traditionally require writing complex Python packages (Pandas, Scikit-Learn, Statsmodels).
Hidden Insights & Patterns
Raw numbers don't tell stories automatically. Identifying key anomalies, linear trends, and correlation matrices takes specialized statistics.
Tool Fragmentation
Juggling separate applications for data cleaning, graphing, AI assistance, and PDF reporting creates friction and data loss.
Black-Box Machine Learning
Standard predictive models give numbers without explanations. DataForge uses SHAP values to explain feature importance clearly.
The DataForge Solution
DataForge orchestrates the complete analytical lifecycle inside a single web workspace. Simply upload your dataset and receive complete end-to-end intelligence:
Automated Data Preprocessing
Imputes missing numerical & categorical values, drops duplicates, and standardizes formats automatically.
Deterministic Statistical Engine
Detects Z-score anomalies, Pearson correlation matrices, trend lines, and contribution breakdowns.
Natural Language AI Querying
Chat with your dataset in plain English powered by Google Gemini 2.5 Flash LLM.
AutoML & SHAP Explainability
Trains predictive models using FLAML and generates interpretable SHAP feature rankings.
How DataForge Works
From raw dataset ingestion to executive PDF reports in 6 automated stages.
Raw Data Ingestion
CSV, Excel & Google Sheets parsing into Pandas DataFrames
Smart Cleaning
Median/Mode imputation, duplicate removal & Z-score bounds
Analytical Engine
Regression trends, Pearson correlations & Holt-Winters forecasting
Executive Output
Interactive Chart.js dashboards & WeasyPrint PDF reports
What Happens After You Upload — DataForge securely ingests, cleans, analyzes, models, and generates interactive reports automatically.
What You Can Do
The DataForge Engine orchestrates 6 interconnected capabilities around your uploaded dataset.
AUTOMATED ANALYTICS PIPELINE
6 Interconnected ModulesSmart Dataset Import
Securely ingest CSV, Excel, or Google Sheets with automatic delimiter detection & schema resolution.
Automated Preprocessing
Automatically detect and impute missing values, remove duplicates, trim whitespace, and clean Z-score outliers.
Deterministic Insights
Statistical algorithms discover linear trends, Pearson correlations, anomaly spikes, and category rankings.
Natural Language Chat
Query your dataset in plain English with Google Gemini LLM to receive instant answers and chart formulas.
AutoML & Explainability
Train FLAML machine learning models with automated feature importance ranking powered by SHAP values.
Executive PDF Exports
Generate comprehensive EDA profiling reports, executive PDF downloads, and shareable dashboards.
Behind the Scenes
Verified, high-performance technologies powering DataForge based on real backend implementation:
Next.js 16 & React 19
Modern responsive web application built with App Router, TypeScript & Tailwind CSS.
FastAPI & Python 3.10+
High-throughput asynchronous Python server executing data cleaning and stats calculations.
Google Gemini 2.5 Flash
Natural language query interpretation generating automated formulas & chart specifications.
FLAML & SHAP Values
Automated model tuning and feature importance scoring for transparent predictions.
Supabase & OAuth 2.0
PostgreSQL database for users and project metadata with Google OAuth authentication.
Chart.js & React-Chartjs-2
Interactive line, bar, scatter, donut, and distribution charts with dark/light themes.
Supabase Storage & Redis
Encrypted file bucket storage for uploaded datasets with Redis task message queues.
WeasyPrint & ydata-profiling
Automated exploratory profiling and executive PDF report synthesis.
What's Next?
Future innovations planned for DataForge to expand automated data engineering:
Advanced Time-Series Forecasting
Expanding Holt-Winters and Prophet models for multi-variable trend forecasting and anomaly prediction.
Multi-Sheet & Table Joining
Automated relational joining and entity synthesis across multiple spreadsheet sheets and connected databases.
Live Cloud Connectors
Direct integrations with Snowflake, BigQuery, AWS S3, and real-time streaming data sources.



