End-to-End AI-Powered Data Analytics Project using Python, Streamlit, LangChain & Groq
Automated Data Analyst is a Streamlit-based application that allows users to upload a dataset, automatically perform exploratory data analysis, and ask questions about the data using natural language.
Instead of manually inspecting rows, calculating statistics, checking missing values, and identifying relationships between columns, the application combines Pandas-based data analysis with an LLM-powered analysis layer to provide a simple conversational experience.
π Live Demo: https://automated-data-analyst.onrender.com/
Note: Since the application is hosted on Render's free tier, the first request may take 30β60 seconds while the server wakes up.
π Dataset Upload
Upload your own datasets directly through the Streamlit interface.
Supported formats:
.csv.xlsx
The application automatically loads the dataset using Pandas.
π Automated Exploratory Data Analysis
Once a dataset is uploaded, the application automatically generates useful dataset information, including:
- Number of rows
- Number of columns
- Numerical columns
- Categorical columns
- Missing values
- Duplicate rows
- Descriptive statistics
- Correlation between numerical variables
This provides an immediate overview of the dataset before asking questions.
π€ Natural Language Data Analysis
Users can ask questions about their dataset using normal language.
For example:
What are the major patterns in this dataset?
Which variables are strongly correlated?
What problems do you see in this dataset?
The application sends the question together with the calculated EDA results to the LLM and generates a concise analytical response.
π§ LLM-Powered Insights
The project uses LangChain + Groq to provide natural-language interpretation of the calculated statistics.
The LLM is instructed to:
- Use the provided EDA results
- Avoid inventing numbers
- Answer the user's question clearly
- Focus on the available dataset information
This separates the numerical computation layer from the natural-language interpretation layer.
π Interactive Streamlit Interface
The application provides a simple interface for:
- Uploading a dataset
- Previewing the data
- Viewing dataset statistics
- Exploring EDA results
- Asking analytical questions
- Receiving AI-generated insights
The application follows a simple and modular pipeline:
ββββββββββββββββββββ
β User Dataset β
β CSV / XLSX β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β Streamlit UI β
β app.py β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β Data Loading β
β data_analysis.pyβ
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β Basic EDA β
β β
β β’ Data Types β
β β’ Missing Values β
β β’ Duplicates β
β β’ Statistics β
β β’ Correlation β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β User Question β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β LangChain + Groq β
β LLM β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β AI Analysis β
β & Insights β
ββββββββββββββββββββ
The project follows a simple principle:
Python calculates the data. The LLM explains the data.
Pandas is responsible for numerical analysis and statistics, while the LLM is used for natural-language interpretation.
| Technology | Purpose |
|---|---|
| Python | Core programming language |
| Streamlit | Web application and UI |
| Pandas | Data loading and analysis |
| NumPy | Numerical operations |
| LangChain | LLM integration |
| Groq | LLM inference |
| OpenPyXL | Excel file processing |
| Plotly | Data visualization support |
| python-dotenv | Environment variable management |
| uv | Python dependency and environment management |
The project currently targets Python 3.14+ and declares its dependencies in pyproject.toml.
Automated-Data-Analyst/
β
βββ app.py # Streamlit application
β
βββ data_analysis.py # Data loading and EDA logic
β
βββ llm.py # Groq LLM configuration and analysis
β
βββ pyproject.toml # Project configuration and dependencies
βββ requirements.txt # Python dependencies
βββ uv.lock # Locked dependency versions
β
βββ .python-version # Python version configuration
βββ .gitignore # Ignored files
β
βββ README.md # Project documentation
Responsible for the application interface and workflow.
It handles:
- Streamlit configuration
- File upload
- Dataset preview
- Dataset metrics
- EDA display
- User questions
- LLM analysis
- Error handling
The current UI exposes metrics such as rows, columns, missing values, and duplicate records.
Contains the core data-analysis functionality.
Responsibilities include:
- Loading CSV/XLSX files
- Detecting numerical columns
- Detecting categorical columns
- Calculating missing values
- Detecting duplicate rows
- Generating descriptive statistics
- Calculating correlations
The EDA layer is implemented using Pandas and NumPy.
Handles the LLM integration.
The application:
- Loads the Groq API key from environment variables.
- Creates a
ChatGroqmodel through LangChain. - Receives the user's question.
- Passes the EDA results to the LLM.
- Generates the final analytical response.
The current implementation uses the qwen/qwen3.8-27b Groq model with deterministic temperature settings.
git clone https://github.com/Swainakash0799/Automated-Data-Analyst.git
cd Automated-Data-Analystuv venvActivate it on Windows:
.venv\Scripts\activateOr on macOS/Linux:
source .venv/bin/activateUsing uv:
uv syncThe repository also includes a uv.lock file for reproducible dependency management.
Create a .env file in the project root:
GROQ_API_KEY=your_groq_api_keyThe application reads the API key using python-dotenv.
β οΈ Never commit your.envfile or expose your API key publicly.
streamlit run app.pyThe application will open in your browser.
Upload a CSV or Excel file.
Example:
retail_customer_shopping_behaviour.csv
The application displays:
Rows
Columns
Missing Values
Duplicates
along with a dataset preview.
View:
- Numerical columns
- Categorical columns
- Missing values
- Descriptive statistics
- Correlation matrix
Example:
What are the main patterns in this dataset?
The application combines the calculated EDA information with your question and generates an easy-to-understand response.
The project currently requires:
| Variable | Description |
|---|---|
GROQ_API_KEY |
API key used to access the Groq LLM |
Example:
GROQ_API_KEY=xxxxxxxxxxxxxxxxYou can ask questions such as:
What are the main characteristics of this dataset?
Are there any missing values or duplicate records?
Which numerical variables appear to be correlated?
What are the most important patterns in this dataset?
What potential data quality issues should I investigate?
The quality of the response depends on the information available in the generated EDA results.
The complete workflow is:
Upload Dataset
β
Load CSV / Excel
β
Inspect Dataset
β
Identify Numerical & Categorical Columns
β
Calculate Missing Values
β
Detect Duplicates
β
Generate Statistics
β
Calculate Correlations
β
User Asks Question
β
EDA Results + Question
β
Groq LLM
β
Natural Language Analysis
This approach keeps data computation deterministic while using the LLM primarily for interpretation.
The application automatically determines the basic structure of the uploaded dataset without requiring the user to manually specify column types.
Core numerical calculations are performed using Pandas rather than relying on the LLM.
For example:
df.describe()and:
df[numerical_columns].corr()This helps keep numerical results grounded in the actual dataset.
The LLM receives the generated EDA information together with the user's question.
The prompt explicitly instructs the model:
Use only the provided EDA results.
Do not invent numbers.
This reduces the risk of unsupported numerical claims.
The current version focuses on:
- CSV/XLSX analysis
- Automated basic EDA
- Data-quality inspection
- Correlation analysis
- Natural-language questions
- LLM-generated explanations
It is intentionally kept lightweight and modular so additional analytical capabilities can be added later.
Planned improvements could include:
- Automated data cleaning
- Interactive Plotly dashboards
- KPI generation
- Advanced visualizations
- Natural-language chart generation
- Statistical hypothesis testing
- Outlier detection
- Time-series analysis
- Forecasting
- Automated business reports
- Downloadable analysis reports
- Conversation history
- Multiple LLM provider support
- Better prompt grounding
- Dataset-aware question answering
- Production deployment
- Automated testing and CI/CD
This project demonstrates practical experience with:
- Python development
- Data analysis with Pandas
- Exploratory Data Analysis
- Data preprocessing
- Data-quality analysis
- Natural Language Processing
- LLM integration
- LangChain
- Groq API
- Prompt engineering
- Streamlit application development
- Environment and dependency management
Traditional data analysis often requires users to:
Load Data
β
Inspect Data
β
Clean Data
β
Calculate Statistics
β
Create Analysis
β
Interpret Results
Automated Data Analyst simplifies this workflow by providing a conversational interface on top of the analytical pipeline:
Upload Data
β
Automatic EDA
β
Ask a Question
β
AI-Assisted Analysis
The goal is not to replace the underlying data-analysis process, but to make it faster and easier to interact with.
Contributions and suggestions are welcome.
If you would like to improve the project:
git clone https://github.com/Swainakash0799/Automated-Data-Analyst.git
cd Automated-Data-AnalystCreate a feature branch:
git checkout -b feature/your-featureMake your changes, test them locally, and open a pull request.
Akash Swain
If you find this project useful, consider giving the repository a β on GitHub.