The Gen AI Powered Data Analyst System is a web-based intelligent analytics platform designed to automate the complete data analysis pipeline using Machine Learning and Large Language Models (LLMs). The system allows users to upload datasets and automatically performs data preprocessing, feature engineering, exploratory data analysis, visualization, and AutoML model selection. It integrates LLMpowered capabilities to generate dataset insights, explain analytical results, interpret model performance, and assist users with natural language understanding of their data. The application provides interactive visualizations, feature importance analysis, prediction results, and downloadable reports through a userfriendlyStreamlit interface. By combining Machine Learning, Generative AI, and intelligent automation, the system reduces manual effort, accelerates data analysis, and supports accurate, explainable, and data-driven decision-making.
Introduction
The Gen AI Powered Data Analyst System is a web-based application designed to automate the end-to-end data analysis process by integrating Machine Learning, AutoML, Large Language Models (LLMs), and interactive visualization. The system addresses the limitations of traditional manual data analysis, which can be time-consuming, error-prone, and dependent on specialized technical expertise. Developed using Streamlit, the application allows users to upload datasets, perform automated preprocessing and feature engineering, visualize data, train multiple machine learning models, and obtain natural-language insights and business recommendations from an LLM.
The literature survey highlights the growing importance of automated machine learning, intelligent preprocessing, dimensionality reduction, interactive visualization, Explainable AI, and Generative AI in modern data analytics. Based on these developments, the proposed system integrates automated handling of missing values and outliers, categorical encoding, datetime extraction, feature selection, scaling, polynomial and interaction features, multicollinearity removal, and optional PCA. These techniques aim to improve data quality, reduce complexity, and enhance machine learning performance.
The proposed methodology begins with dataset upload and automated data preprocessing, followed by intelligent feature engineering and exploratory data analysis. An AutoML pipeline identifies whether the problem is classification or regression, trains several algorithms—including Logistic Regression, Decision Tree, Random Forest, SVM, KNN, and Linear Regression—and selects the best model according to suitable evaluation metrics. The system also generates correlation heatmaps, distribution plots, feature-importance charts, confusion matrices, ROC curves, and model-comparison graphs to support interpretation.
A major feature of the system is the integration of LLMs, which generate dataset summaries, explain preprocessing steps and model performance, answer user questions, and provide business-oriented recommendations in natural language. The overall architecture consists of data processing, AutoML, LLM, and presentation layers, providing a unified workflow from raw data to predictions and actionable insights.
The system outputs include an interactive home dashboard, exploratory data analysis, outlier detection and treatment, feature engineering and selection, automated model training, prediction visualization, and comparative model-performance analysis. Users can also download processed datasets, trained models, and prediction reports.
Conclusion
The proposed Gen AI Powered Data Analyst System provides an intelligent and efficient solution for data analysis using machine learning techniques. It automates data preprocessing, feature engineering, visualization, and model selection through AutoML, reducing manual effort and improving accuracy. Interactive visualizations enable better understanding of data patterns and model performance. The system offers a user-friendly Streamlit interface for seamless analysis and decision-making. Its modular architecture ensures scalability and flexibility for various datasets. Future enhancements can include real-time data integration, LLM-based insights, and advanced explainable AI techniques.
References
[1] F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011. [Online]. Available: https://scikit-learn.org
[2] W. McKinney, “Data Structures for Statistical Computing in Python,” Proceedings of the Python in Science Conference, 2010. [Online]. Available: https://pandas.pydata.org
[3] C. R. Harris et al., “Array Programming with NumPy,” Nature, vol. 585, pp. 357–362, 2020. [Online]. Available: https://numpy.org
[4] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” Proceedings of the ACM SIGKDD Conference, 2016. [Online]. Available: https://xgboost.readthedocs.io
[5] L. Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001. [Online]. Available: https://link.springer.com/article/10.1023/A:1010933404324
[6] H2O.ai, “H2O AutoML Documentation,” 2025. [Online]. Available: https://docs.h2o.ai
[7] J. H. Friedman, “Greedy Function Approximation: A Gradient Boosting Machine,” The Annals of Statistics, 2001. [Online]. Available: https://projecteuclid.org
[8] Microsoft, “Automated Machine Learning (AutoML),” Azure Machine Learning Documentation, 2025. [Online]. Available: https://learn.microsoft.com/azure/machine-learning
[9] Streamlit Inc., “Streamlit Documentation,” 2025. [Online]. Available: https://streamlit.io
[10] J. VanderPlas, Python Data Science Handbook, O\'Reilly Media, 2016. [Online]. Available: https://jakevdp.github.io/PythonDataScienceHandbook
[11] I. T. Jolliffe, Principal Component Analysis, 2nd ed., Springer, 2002. [Online]. Available: https://link.springer.com/book/10.1007/b98835
[12] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning, Springer, 2009. [Online]. Available: https://hastie.su.domains/ElemStatLearn
[13] M. Kuhn and K. Johnson, Feature Engineering and Selection: A Practical Approach for Predictive Models, CRC Press, 2019. [Online]. Available: https://bookdown.org/max/FES
[14] J. Brownlee, “Feature Engineering for Machine Learning,” Machine Learning Mastery, 2023. [Online]. Available: https://machinelearningmastery.com
[15] [15] Plotly Technologies Inc., “Plotly Python Graphing Library,” 2025. [Online]. Available: https://plotly.com/python
[16] [16] S. Raschka and V. Mirjalili, Python Machine Learning, 3rd ed., Packt Publishing, 2019. [Online]. Available: https://www.packtpub.com
[17] [17] M. T. Ribeiro, S. Singh, and C. Guestrin, “Why Should I Trust You? Explaining the Predictions of Any Classifier,” Proceedings of the ACM SIGKDD Conference, 2016. [Online]. Available:https://arxiv.org/abs/1602.04938
[18] A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras& TensorFlow, 3rd ed., O\'Reilly Media, 2022. [Online]. Available: https://www.oreilly.com/library/view/hands-onmachinelearning/9781098125967/