Understanding multilingual social media behavior over time is hard because online archives hold large and varied informal posts and because people talk and interact in ways that keep changing. The paper introduces an Agent-Based Framework for Longitudinal Multilingual Social Media Behavior Analysis that uses XLM?RoBERTa based Multi?Task Learning. The framework uses software agents to collect archives pull out data, clean multilingual text analyze behavior organize data over time create profiles and show results. Social media records that include posts, comments, messages, replies and activity data are looked at in 3?, 6? and 12?month windows to find patterns and trends. XLM?RoBERTa is used with heads for sentiment, emotion and hate speech analysis. Experiments show 79.96?% accuracy and 77.49?% Macro?F1 for emotion , 69.15?% accuracy and 69.27?% Macro?F1 for sentiment and 83.07?% accuracy and 82.68?% Macro?F1, for hate speech. The framework gives dashboards for statistics, time trends, engagement, language and interaction. Results show that combining agent?based processing, multilingual transformer learning and longitudinal analysis can give a picture of multilingual social media behavior.
Introduction
The text presents a multilingual, time-aware framework for analyzing social media behavior using XLM-RoBERTa, Multi-Task Learning (MTL), and software agents.
Problem: Social media contains huge amounts of multilingual, informal, and code-mixed content. User behavior also changes over time, making it difficult to understand behavior by analyzing individual posts alone.
Research gap: Existing studies mainly perform separate tasks such as sentiment or emotion analysis and often focus on individual posts. They rarely combine multilingual analysis, multiple behavioral characteristics, and long-term behavior patterns in one system.
Proposed solution: The paper introduces an agent-based framework that analyzes social media activity over 3, 6, or 12 months to create comprehensive user behavior profiles.
Core technologies: The system combines XLM-RoBERTa, multilingual Natural Language Processing, Multi-Task Learning, temporal analysis, and autonomous software agents.
Behavioral analysis: The shared model simultaneously predicts sentiment, emotion, toxicity, distress, conflict, and engagement, allowing information to be shared across related tasks and improving efficiency.
System architecture: Different agents perform specific functions, including:
Data Acquisition Agent – collects social media data.
Data Structuring Agent – organizes information into a common format.
Multilingual Preprocessing Agent – identifies languages, cleans text, and handles English, Hindi, Telugu, and code-mixed content.
Behavior Intelligence Agent – uses XLM-RoBERTa and MTL for behavioral predictions.
Temporal Intelligence Agent – analyzes behavioral changes over time.
Behavioral Profiling Agent – creates a long-term profile of each user.
Visualization and Analytics Agent – presents results through dashboards and reports.
Datasets: The framework uses datasets such as GoEmotions, Jigsaw Multilingual Toxic Comment Classification, stress-related Reddit data, SentiMix, and Dravidian-CodeMix to support emotion, toxicity, stress, sentiment, and multilingual analysis.
Data preprocessing: Social media content is cleaned by handling HTML, links, hashtags, mentions, emojis, slang, abbreviations, spelling variations, and informal language.
Multilingual processing: SentencePiece tokenization and XLM-RoBERTa allow the framework to process multiple languages and code-mixed text without requiring completely separate models for each language.
Main contribution: Instead of providing a snapshot of individual posts, the framework builds a long-term and multidimensional picture of social media behavior, capturing communication patterns, emotions, engagement, language use, conflict, and behavioral evolution.
Conclusion
In this paper, we have suggested an agent-based pipeline for longitudinal social media behavior analysis based on XLM-RoBERTa transformer learning model. Data acquisition, archive processing, data structuring, multilingual pre-processing, behavior analysis, temporal aggregation, behavioral profiling and visualization modules are all included in the agent-based framework designed in the paper.The experimental evaluation has demonstrated the applicability of XLM-RoBERTa transformer model for important social media language analysis tasks. The model was able to achieve 69.15% accuracy and 69.27% Macro F1-score in sentiment analysis, whereas emotion recognition and hate speech detection models have shown the applicability of transformers for the purpose of social media text analysis. In addition, our framework allows for performing temporal aggregation using 3, 6 and 12-month analysis windows, which would allow for detecting changes and trends in behavioral indicators.The proposed architecture suggests a modular structure for incorporating various behavioral analysis tasks and representing their outputs in interactive dashboards, statistical reports, trend visualizations and longitudinal profiles.
Future works will include further expansion of the proposed framework in different ways. The first one includes integration of the individually trained sentiment, emotion and hate speech models into a multi-task framework based on a single XLM-RoBERTa transformer encoder with different prediction heads for each task. The comparative experiments involving the individual and multi-task frameworks could reveal advantages of the shared representation learning.
Secondly, the multilingual and code-mixed capabilities of the framework can be validated with multilingual and code-mixed datasets, which include regional languages and code-mixed samples. Other behavioral tasks, such as toxicity, distress, sarcasm, expression intensity, and engagement can be introduced after validation of the dataset.
Thirdly, the future works will involve exploration of different methods for improving the model\'s robustness and reproducibility, such as hyper-parameter optimization, class balancing, learning rate schedule, data augmentation and use of multiple random seeds. Advanced temporal modeling can also be considered in order to account for behavioral dynamics and long-term trends not limited by fixed 3-, 6-, and 12-month aggregation windows.
References
[1] Amani N. Alansari, Abdallah Alwawi, Amani Salim, et al., \"A systematic review of social media impact,\" Asian Journal of Psychiatry, vol. 115 2026.
[2] Prema Nedungadi, G. Veena, Kai-Yu Tang, et al., \"AI Techniques and Applications forOnline Social Networks and Media: Insights From BERTopic Modeling,\" IEEE Access, 2025.
[3] Darji, F. Kheni, D. Chodvadia, P. Goel, D. Garg and Patel, “Enhancing Financial Risk Analysis using RAG-based Large Language Models,” in 2024 3rd International Conference on Automation, Computing and Renewable Systems (ICACRS), Pudukkottai, India, 2024, pp. 754–760.
[4] Suman Kumari, \"EXAMINING THE IMPACT OF SOCIAL MEDIA ON MENTAL HEALTH OF COLLEGE STUDENTS,\" International Journal of Management Research and Analysis (IJMRA), vol. 15, iss. 5, 2025.
[5] N. Chinaksorn and D. Wanvarie, “LLM-RAG for Financial Question Answering: A Case Study from SET50,” in 2025 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Fukuoka, Japan, 2025, pp. 952–957.
[6] Nadine Agyapong-Opoku, Felix Agyapong-Opoku, Andrew J. Greenshaw, \"Effects of Social Media Use on Youth and Adolescent Mental Health: A Scoping Review of Reviews,\" Behavioral Sciences, vol. 15, iss. 5, 574, 2025.
[7] Sidra Hameed, Muhammad Nauman, Nadeem Akhtar, et al., \"Explainable AI-driven depression detection from social media using natural language processing and black box machine learning models,\" Frontiers in Artificial Intelligence, vol. 8, 1627078, 2025.
[8] Nisha P. Shetty, Yashraj Singh, Veeraj Hegde, et al., \"Exploring emotional patterns in social media through NLP models to unravel mental health insights,\" Healthcare Technology Letters, vol. 12, e12096, 2025.
[9] Shiwen Zhou, Masnizah Mohd, \"Mental Health Safety and Depression Detection in Social Media Text Data: A Classification Approach Based on a Deep Learning Model,\" IEEE Access, 2025.
[10] Qasim Bin Saeed, YoungJin Cha, \"Multi-modal deep-attention-BiLSTM based early detection of mental health issues using social media posts,\" Scientific Reports (Nature Portfolio), 2025.
[11] Noemí Merayo, Alba Ayuso-Lanchares, Clara González-Sanguino, \"Revealing Emotional Insights From Mental Health Discussions on Instagram and TikTok Using BERT Models,\" IEEE, 2025.
[12] Jessica L. Hamilton, Srushti Untawale, Maya N. Dalack, et al., \"Self-Harm Contenton Social Media and Proximal Risk,\" Journal of the American Academy of Child & Adolescent Psychiatry, 2025.
[13] Masab A. Mansoor, Kashif H. Ansari, \"Early Detection of Mental Health Crises through Artificial-Intelligence-Powered Social Media Analysis: A Prospective Observational Study,\" Journal of Personalized Medicine, vol. 14, 958, 2024.
[14] Deepanjal Sood, Ankit Anupam Rout, Ishmeet Singh, \"Multimodal AI Framework forEarly Detection of Mental Health Disorders via Social Media Analysis,\" Internationa lJournal for Innovative Research in Technology (IJIRT), vol. 12, iss. 3, 2025.