Choosing the right research topic can make a big difference when you’re getting started with Data Science in high school. A topic that is too broad can quickly become overwhelming, making it hard to figure out where to start, what questions to ask, and what data to analyze. On the other hand, a focused, specific idea makes it easier to collect data, build a clear research question, and put together a stronger final paper, presentation, or project. A well-defined topic also helps you get more meaningful feedback from teachers, mentors, and professors. If you’re planning to include research projects in your college applications, a clear and thoughtful Data Science topic can also demonstrate your problem-solving skills and technical interests.
The good news is that you do not need advanced college-level knowledge to conduct impactful data science research. You might use tools such as Python or R to analyze datasets, data visualization libraries to present findings, simple machine learning models to make predictions, and statistical methods to study patterns and trends. Picking the right research topic can make your Data Science project much easier and a lot more interesting.
How do you pick a good Data Science topic?
A good Data Science topic should be specific, researchable, and interesting to you. For example, “Social media and mental health” is too broad, while “Analyzing whether Instagram screen time affects sleep patterns among teenagers” is much more focused and workable. Before committing to a topic, you should check whether you can actually find datasets, articles, or sources to support your research. Choosing a topic connected to something you already enjoy, whether that is sports, social media, gaming, healthcare, music, or another area of interest, can make the research process feel much less overwhelming. It also helps if your research topic connects to a real-world problem or question that can be investigated using data.
To help you get started, we’ve put together this list of 25 Data Science research topics covering a range of skill levels, subfields, and project styles. Some topics are beginner-friendly and focus on visualization or basic statistics, while others introduce concepts such as machine learning, predictive modeling, and larger datasets.
If you’re looking for data science research programs, you can find them here. You can also check out AI research topics for high school students here.
What is Data Science research?
Data Science research at the high school level involves the use of data to answer questions, study patterns, and draw meaningful conclusions about real-world topics. By collecting and interpreting data, you can investigate everything from social media habits to climate trends and public health issues. You can use Python or R to clean datasets, create graphs, run basic statistical tests, or build simple machine learning models. You can also work with publicly available datasets from sources such as Kaggle, government databases, and research organizations, making it possible to conduct meaningful research without collecting all of your own data. A typical project starts with a clear research question, such as “Can weather data help predict air quality?” You then collect data, analyze it, visualize the results with charts or graphs, and explain what your findings reveal. Basic coding and statistics skills can come in handy, but plenty of beginner-friendly resources are available online to help you get started. Some common projects focus on topics such as social media use, sports analytics, climate trends, public health, and even music recommendations.
Here are 25 Data Science research topics for high school students:
Environmental and Climate Data Science
This field focuses on using data to better understand environmental changes, pollution, weather patterns, and natural disasters. As a student researcher, you’ll work with climate records, satellite imagery, sensor data, and geographic information to investigate real-world environmental problems. If you’re interested in learning about sustainability, climate science, or mapping trends across regions, this field gives you many beginner-friendly research opportunities.
Key Takeaways
- 25 data science research topics span 10 subfields, from environmental and climate science to sports analytics, giving students options across nearly any area of personal interest.
- A good topic needs to be specific, researchable, and personally interesting. “Social media and mental health” is too broad, while “Analyzing whether Instagram screen time affects sleep patterns among teenagers” is focused enough to build a clear methodology around.
- Many topics rely entirely on publicly available datasets from sources like Kaggle, the EPA, the CDC, the Census Bureau, and NASA, meaning students can conduct meaningful research without collecting their own data or needing institutional access.
- Difficulty levels vary considerably within this list. Beginner topics like studying the relationship between sleep and academic performance or analyzing income inequality through Census data require only spreadsheets and basic statistics, while advanced topics like predicting disease outbreaks or measuring misinformation spread during disasters require machine learning models and more sophisticated statistical tools.
- Python and R are the most commonly used tools across these topics, with specific libraries like Pandas, Scikit-learn, NLTK, and Matplotlib recurring throughout projects in natural language processing, predictive modeling, and data visualization.
- Several topics connect directly to current events and active public debates. Identifying bias in AI language models and analyzing social media sentiment during elections both let students engage with questions that are still being actively researched and discussed.
- Data science research is a strong way to build technical and analytical skills, but it is not the only path for a student interested in independent research in other fields. Those who want to pursue a research project outside data science might consider Horizon’s Research Seminars and Labs, a mentored virtual research program with 600+ specializations across subjects like political theory, biology, and machine learning.
1. Predicting Air Quality Using Traffic and Weather Data
Sub-fields: Environmental data science, atmospheric science, urban transportation research, spatial statistics, predictive modeling, environmental health informatics
Difficulty Level: Beginner
Research Question: How accurately can traffic volume and temperature predict daily air quality in urban areas?
Air quality can significantly vary throughout the year due to changes in weather conditions and seasonal factors. In this project, you can investigate how variables such as vehicle traffic and weather conditions affect pollution levels in a city or region. This topic matters because poor air quality is linked to respiratory illnesses such as asthma and other health problems. To conduct your research, you can use public datasets from the EPA (Environmental Protection Agency), local traffic records, and weather databases, and then analyze them with Python and spreadsheets to identify trends and compare pollution levels across different time periods. This topic works well if you are interested in climate issues, public health, or beginner-level coding and statistics.
2. Predicting Wildfire Risk Using Climate Data
Sub-fields: Environmental data science, climatology, remote sensing, geospatial analysis, ecology, predictive modeling, disaster risk science
Difficulty Level: Intermediate
Research Question: Which climate variables are most strongly associated with wildfire outbreaks?
Wildfires have become more frequent in places such as California, making wildfire prediction an important area of research. In this project, you examine how rising temperatures, drought conditions, and wind patterns contribute to wildfire risk. To conduct your research, you can analyze satellite datasets such as NASA MODIS fire imagery and Landsat imagery, weather records from NOAA or local meteorological agencies, and wildfire databases using Python or GIS software. This topic is a good choice if you enjoy environmental science, climate studies, or working with large datasets.
3. Using Satellite Data to Study Deforestation
Sub-fields: Remote sensing, geospatial analysis, environmental data science, land cover science, conservation biology, climate science
Difficulty Level: Intermediate
Research Question: Which regions have experienced the fastest deforestation rates in the last decade?
Deforestation affects biodiversity, climate systems, and indigenous communities around the world. This project allows you to analyze satellite imagery to track forest loss and land-use changes over time. To conduct your research, you can use NASA or Global Forest Watch datasets along with GIS software and image analysis tools. This topic is a good choice if you are interested in environmental science or geographic data analysis.
4. Forecasting Water Usage During Drought Conditions
Sub-fields: Hydrology, climatology, environmental data science, time series forecasting, water resource management, geospatial analysis
Difficulty Level: Intermediate
Research Question: How accurately can weather conditions predict household water consumption during droughts?
For this topic, you explore how communities change their water usage during droughts, heat waves, and periods of water use restrictions. Water shortages are becoming common in many regions because of climate change and population growth. To conduct your research, you can analyze municipal water records from local city water boards, weather datasets from NOAA, NASA, or national meteorological departments. You can also use Python libraries such as Pandas to organize data and Scikit-learn for basic forecasting models. This project is a good fit if you are interested in sustainability, environmental science, or applied statistics.
Health and Public Health Analytics
This field involves using medical, behavioral, and population-level data to understand health outcomes and improve healthcare systems. You might analyze data from sources such as survey results, patient records, or disease surveillance systems to identify risk factors and trends in disease outbreaks. This field is a great option if you are interested in learning about medicine, psychology, biology, or healthcare policy.
5. Studying the Relationship Between Sleep and Academic Performance
Sub-fields: Behavioral data science, educational data science, cognitive psychology, survey methodology, biostatistics
Difficulty Level: Beginner
Research Question: Does average sleep duration affect student grades?
Sleep directly affects concentration and memory, yet many teenagers struggle to get enough rest because of their busy schedules. For this topic, you can analyze whether students who sleep longer tend to perform better in school. To conduct your research, you can collect survey data, use statistical analysis in spreadsheets or Python, and compare sleep patterns with GPA or test scores. This topic works well if you’re interested in psychology, education, or beginner-level statistics.
6. Predicting Disease Outbreaks from Public Health Data
Sub-fields: Epidemiology, public health informatics, biostatistics, predictive modeling, spatial epidemiology, computational biology
Difficulty Level: Advanced
Research Question: Can population density and vaccination rates predict outbreak severity?
For this topic, you’ll examine how demographic and healthcare variables influence the spread of infectious diseases. Public health agencies increasingly rely on predictive models after the COVID-19 pandemic exposed weaknesses in outbreak forecasting. You can use the Centers for Disease Control and Prevention (CDC) datasets, such as flu surveillance reports and chronic disease statistics, as well as statistical modeling and Python visualization tools such as Matplotlib or Plotly, to conduct your research. This topic is a great choice if you’re interested in biology, healthcare, or advanced statistics.
7. Predicting Hospital Readmission Rates
Sub-fields: Healthcare analytics, clinical informatics, biostatistics, predictive modeling, health services research
Difficulty Level: Advanced
Research Question: Which patient factors are most strongly associated with hospital readmissions?
For this topic, you study how hospitals use patient data to reduce repeat admissions and improve care quality. Hospital readmissions are expensive for healthcare systems and often linked to gaps in follow-up treatment. You can analyze healthcare datasets, including hospital admission records and patient outcome databases, using machine learning models such as Scikit-learn in Python and statistical software such as SPSS, R, or Python libraries such as Pandas and Statsmodels. This project works well if you are interested in medicine, public health, or advanced data science.
Natural Language Processing and Computational Linguistics
This area focuses on analyzing written and spoken language using machine learning and text analysis techniques. You’ll work with articles, reviews, conversations, or social media posts to study sentiment, bias, misinformation, or communication patterns. This field is a good fit if you enjoy learning about language, artificial intelligence, social media, or computational analysis.
8. Analyzing Social Media Sentiment During Elections
Sub-fields: Computational social science, political science, natural language processing, electoral studies, sentiment analysis
Difficulty Level: Intermediate
Research Question: How does public sentiment on social media change before and after major election debates?
For this topic, you can explore how people react online during election cycles by analyzing posts from platforms such as X or Reddit. This topic is particularly interesting because online discussions influence political opinions and spread misinformation quickly, making social media an important part of modern elections. To conduct your research, you can collect text data using APIs (Application Programming Interface), analyze it using Python libraries, and compare trends around major events. This project is a good fit if you enjoy politics, media analysis, or language-based data science and have some basic coding experience.
9. Detecting Cyberbullying in Online Conversations
Sub-fields: Natural language processing, computational social science, human-computer interaction, adolescent psychology, online safety research
Difficulty Level: Intermediate
Research Question: Which text patterns are most associated with cyberbullying behavior online?
For this topic, you can explore how language analysis and machine learning can be used to automatically identify harmful online interactions. Cyberbullying has become a growing concern for schools and social media platforms, making it an important area of research. You can analyze public conversation datasets, such as Twitter hate-speech datasets, Kaggle cyberbullying datasets, Reddit comment collections, or online chat datasets, and build simple text classification models using Python to identify harmful and abusive language. This project is a great option if you care about online safety, psychology, or machine learning.
10. Identifying Bias in AI Language Models
Sub-fields: AI ethics, natural language processing, fairness and accountability in ML, computational linguistics, social computing
Difficulty Level: Advanced
Research Question: How do AI language models respond differently to prompts involving gender or ethnicity?
For this topic, you can investigate whether AI systems produce biased or unequal responses across different groups. Bias in artificial intelligence is widely debated because these systems are increasingly used in areas such as hiring, healthcare, and education. To conduct your research, you can compare outputs from different language models, such as ChatGPT, Gemini, Claude, or open-source models including Llama, categorize response patterns, and analyze results statistically using Python libraries such as Pandas and NumPy. This topic is a great choice if you’re interested in ethics, computer science, or AI research.
11. Analyzing Climate Change Discussions on Social Media
Sub-fields: Computational social science, environmental communication, natural language processing, science communication, media studies
Difficulty Level: Intermediate
Research Question: How does public discussion about climate change shift after extreme weather events?
In this project, you can find out how people react online after heat waves, floods, or hurricanes linked to climate change. Public opinion about climate change often changes after visible environmental events. To conduct your research, you can collect social media posts from platforms such as X, Reddit, or YouTube comments, analyze keywords and sentiment using Python libraries such as NLTK, TextBlob, or VADER, and compare trends over time. This project works well if you’re interested in environmental policy, communication, or text analysis.
Financial and Economic Analytics
This field uses statistics and data analysis to understand financial markets, investments, income trends, and consumer behavior. As a student researcher, you’ll work with economic indicators, past market data, and financial transactions to identify patterns, make predictions, and study how financial systems operate. This field is a strong choice if you’re interested in finance, economics, business, or applied mathematics.
12. Forecasting Stock Prices Using Historical Market Data
Sub-fields: Financial data science, quantitative finance, time series analysis, econometrics, predictive modeling
Difficulty Level: Intermediate
Research Question: How accurately can historical price trends predict short-term stock movements?
For this topic, you can study whether past stock market behavior helps forecast future price changes. Financial analysts use similar models every day, but markets remain difficult to predict because news and investor emotions affect prices. You can use Yahoo Finance datasets, Python libraries such as Pandas, and time-series analysis methods to identify trends and test whether historical data can help predict future market movements. This project is a good fit if you are interested in finance, economics, or statistical modeling.
13. Detecting Trends in Cryptocurrency Markets
Sub-fields: Financial data science, quantitative finance, time series analysis, network science, behavioral economics
Difficulty Level: Intermediate
Research Question: Which market indicators best predict short-term cryptocurrency price changes?
Cryptocurrency markets are known for rapid price swings and unpredictable investor behavior. For this topic, you can explore how digital currencies respond to factors such as market sentiment, regulation, and social media activity. You can use trading datasets such as Kaggle cryptocurrency datasets, Reddit discussions, and Python statistical modeling tools such as Pandas or NumPy to organise data and identify trends. This topic is a great option if you’re interested in finance, economics, or emerging technology trends.
14. Analyzing Income Inequality Through Census Data
Sub-fields: Economic data science, econometrics, social statistics, demography, public policy analysis
Difficulty Level: Beginner
Research Question: How does income inequality differ between urban and rural regions?
For this project, you can study how income levels differ by geography, education, and employment type. Income inequality affects housing, education access, and healthcare outcomes in many communities. You can use Census Bureau datasets, spreadsheets, and data visualization software such as Tableau and Power BI, or Python libraries such as Matplotlib and Plotly, to compare income trends and identify income inequality patterns across different patterns and geographic areas. This project works well if you are interested in economics, sociology, or public policy.
Social Science and Public Policy Analytics
This area focuses on using data to better understand the behavior of people, communities, and governments. You’ll analyze demographic records, surveys, election data, or online discussions to study social issues and the impact of policy on different groups and communities. This field is a good choice if you are interested in politics, sociology, education, or community-focused research, as it connects data science with real-world decision-making.
15. Measuring the Spread of Misinformation During Natural Disasters
Sub-fields: Computational social science, crisis informatics, natural language processing, network science, disaster communication
Difficulty Level: Advanced
Research Question: How quickly does misinformation spread compared to verified information during disasters?
Emergency agencies struggle with misinformation because it tends to affect evacuation decisions and public safety. For this topic, you’ll investigate how false information circulates online during natural disasters such as hurricanes, earthquakes, or wildfires. You can track social media posts, analyze sharing networks, and compare verified information versus false claims using Python. This topic is a great option if you enjoy social media analysis, network science, or public policy research.
16. Predicting Voter Turnout Using Demographic Data
Sub-fields: Political data science, electoral studies, demography, computational social science, survey methodology
Difficulty Level: Intermediate
Research Question: Which demographic variables are most associated with voter turnout rates?
For this topic, you can study how factors such as age, education, and income influence election participation. Governments and advocacy groups use similar research to improve civic engagement. To conduct your research, you can analyze election datasets, such as state- or county-level voter turnout records, Census data, and regression models, in Python or Excel. This project works well if you are interested in politics, sociology, or statistical analysis.
Urban, Transportation, and Geographic Analytics
This field focuses on understanding how cities function through the analysis of transportation systems, infrastructure, geographic trends, and population movement. As a student researcher, you’ll use maps, GPS data, public transit records, and geographic datasets to explore topics such as traffic congestion, urban growth, and infrastructure planning. This field is a strong fit if you’re interested in city planning, engineering, geography, or visual data analysis.
17. Predicting Traffic Congestion Patterns in Cities
Sub-fields: Urban computing, transportation engineering, spatial statistics, time series forecasting, geospatial analysis
Difficulty Level: Beginner
Research Question: Which factors most strongly influence rush hour traffic congestion?
For this topic, you can investigate how factors such as weather, time of day, and public events affect traffic flow in urban areas. Traffic congestion costs cities time and fuel, and also increases pollution levels. To conduct your research, you can analyze transportation datasets such as city traffic volume records or public transit schedules, GPS records, and weather data using spreadsheets or Python to compare traffic conditions across different times and locations. This topic is a good match if you are interested in city planning, transportation, or beginner data visualization.
18. Analyzing Public Transportation Efficiency
Sub-fields: Urban computing, transportation engineering, operations research, geospatial analysis, public policy analysis
Difficulty Level: Intermediate
Research Question: What factors contribute most to delays in public transportation systems?
Cities increasingly depend on data analysis to improve transportation reliability and reduce commuter frustration. In this project, you can explore how schedules, weather, and passenger volume affect train or bus delays. To conduct your research, you can analyze transit datasets such as New York City Subway ridership data or Chicago Transit Authority bus and train records, and GPS records such as Google Maps traffic data. Using Python and visualization tools such as Matplotlib, you can identify trends and compare delays across different routes and time periods. This topic is a good option if you’re interested in engineering, urban planning, or operations research.
19. Mapping Food Deserts in Urban Communities
Sub-fields: Geospatial analysis, urban sociology, public health informatics, social equity research, environmental justice
Difficulty Level: Beginner
Research Question: Which neighborhoods have the lowest access to affordable grocery stores?
For this topic, you can explore how access to healthy food varies across communities and how it is connected to factors such as income and transportation. Food deserts, i.e., areas where residents have limited access to affordable and nutritious food, are linked to nutrition and health problems in many cities. You can use Census data, GIS mapping tools such as ArcGIS and Google Earth, and local business records to identify areas with limited food access. This topic is a great option if you’re interested in public policy, geography, or community-focused research.
Media, Culture, and Entertainment Analytics
This category looks at how data can reveal trends in entertainment, communication, and popular culture. If you enjoy pop culture, storytelling, or digital platforms, this field offers a wide range of creative research opportunities with accessible datasets. You might get to study movies, music, online communities and audience behavior to understand how media affects public interest and opinions.
20. Predicting Movie Box Office Success from Online Reviews
Sub-fields: Sentiment analysis, natural language processing, cultural analytics, predictive modeling, media economics
Difficulty Level: Beginner
Research Question: Can online review scores and trailer views predict opening weekend revenue?
Film studios rely heavily on digital marketing, but predicting a film’s financial success is still difficult because audience preferences and entertainment trends can shift quickly. For this topic, you can investigate whether audience reactions before a movie’s release influence its ticket sales. You can gather IMDb ratings, YouTube trailer statistics, and box office numbers to conduct your research, and then use regression analysis in Python or Excel to compare audience metrics with ticket sales. This topic is a good choice if you’re interested in entertainment, marketing, or beginner-friendly data analysis projects.
21. Identifying Music Trends Through Streaming Data
Sub-fields: Cultural analytics, music information retrieval, behavioral data science, time series analysis, recommender systems
Difficulty Level: Beginner
Research Question: How do seasonal patterns influence music streaming popularity across genres?
Music streaming platforms constantly collect user behavior data, making them valuable sources for understanding cultural trends and shifts in audience preferences. For this topic, you can study how listening habits change during holidays, school breaks and viral online trends. You can use Spotify datasets such as Spotify Charts or Spotify API data to conduct your research. Using data visualization tools such as Tableau and Power BI, Python libraries, and statistical analysis software including Python, R, SPSS, or Excel, you can identify trends and compare listening behavior. This topic works well if you enjoy music, pop culture, or beginner-friendly data science projects.
22. Examining Gender Representation in Film Dialogue
Sub-fields: Computational humanities, gender studies, cultural analytics, natural language processing, media studies
Difficulty Level: Intermediate
Research Question: Do male and female characters receive equal speaking time in popular films?
For this topic, you analyze movie scripts to measure how dialogue is distributed among characters. Representation in media remains widely debated, especially in blockbuster films and award-winning movies. You can use screenplay databases such as IMSDb and SimplyScripts to conduct your research, analyze movie scripts using text analysis tools such as NLTK, spaCy, or TextBlob, and use Python libraries to compare speaking patterns across different groups and time patterns. This project is a good match if you’re interested in media studies, gender studies, or computational text analysis.
Sports and Performance Analytics
This field uses statistics and performance data to evaluate athletes, teams, and game strategies. As a student researcher, you’ll examine player efficiency, scoring trends, or team performance to understand what contributes to success in sports. This field is a good choice if you enjoy athletics, math, or competitive strategy.
23. Analyzing Sports Performance Using Player Statistics
Sub-fields: Sports analytics, performance science, biostatistics, predictive modeling, exploratory data analysis
Difficulty Level: Beginner
Research Question: Which player statistics best predict team success in basketball?
Sports organizations increasingly rely on analytics for recruiting and strategy decisions. For this topic, you can investigate how individual performance contributes to overall team results in professional sports. You can use NBA or other sports datasets, statistical analysis methods, and visualization tools such as NumPy, Scikit-learn, or Matplotlib. This project is great if you’re interested in athletics, math, or applied statistics.
24. Predicting Athlete Injuries Using Training and Performance Data
Sub-fields: Sports analytics, biomechanics, predictive modeling, health informatics, performance science, biostatistics, wearable sensor data analysis
Difficulty Level: Intermediate
Research Question: Which training variables are most strongly associated with injury risk in athletes?
Professional teams increasingly use data science to reduce injuries because overtraining can affect both athlete health and team performance. For this topic, you can analyze how workload, recovery time, and performance statistics relate to sports injuries. You can work with player-tracking datasets, such as GPS-based athlete-tracking systems, wearable fitness device data, fitness records, and statistical models, using Python or spreadsheets to identify injury patterns. This topic works well if you are interested in athletics, health science, statistics, or performance analysis.
Educational Analytics
Educational analytics focuses on understanding how students learn, participate, and perform in academic settings. As a student researcher, you might analyze grades, surveys, engagement metrics and admissions data to investigate topics such as academic performance and learning behaviors, and identify patterns in learning and school systems. If you are interested in education, psychology or social science research, this field offers practical and accessible project opportunities.
25. Measuring the Impact of Online Learning on Student Engagement
Sub-fields: Educational data science, learning analytics, human-computer interaction, educational psychology, survey methodology
Difficulty Level: Beginner
Research Question: How does participation in online classes affect assignment completion rates?
Schools continue to debate the long-term effectiveness of online education, especially since the COVID-19 pandemic increased remote learning. For this topic, you’ll study how virtual learning environments influence student behavior and academic participation. You can use survey responses collected through tools such as Google Forms or SurveyMonkey, along with learning management system data such as attendance records or assignment submission rates, and statistical analysis tools including Google Sheets or Python libraries to examine different aspects of online learning. This topic is a great match if you’re interested in learning about education, psychology, or beginner-level research methods.
26. Predicting College Admissions Outcomes from Applicant Data
Sub-fields: Educational data science, predictive modeling, fairness and accountability in ML, higher education research, social statistics
Difficulty Level: Intermediate
Research Question: Which application factors most strongly influence college admissions decisions?
College admissions are competitive and sometimes controversial because schools prioritize different criteria. For this project, you’ll study how GPA, extracurricular activities, and standardized test scores influence admissions outcomes. You can analyze publicly available admissions datasets, such as Kaggle college admissions datasets or Common Data Set reports published by universities, by using regression or classification models. This project works well if you are interested in education policy or predictive analytics.
Internet of Things and Smart Technology
This area focuses on data collected from connected devices such as smart home systems, sensors, and wearable technology. As a student researcher, you’ll analyze real-time data streams to improve efficiency, monitor behavior, and predict usage patterns. This field is a good fit if you are interested in engineering, automation, and technology-driven research.
27. Predicting Energy Consumption in Smart Homes
Sub-fields: Energy informatics, IoT data science, time series forecasting, environmental data science, smart systems engineering
Difficulty Level: Intermediate
Research Question: Which household activities are most associated with spikes in energy use?
Energy companies are always interested in reducing waste and improving efficiency through data-driven systems. For this topic, you can analyze how smart devices track electricity use in homes throughout the day. You can use smart home datasets such as smart thermostat records, smart meter electricity usage data, sensor data, and time-series analysis in Python to identify trends in energy consumption and the factors that affect household energy use. This topic is a strong fit if you enjoy technology, sustainability, or engineering-related projects.
28. Real-Time Flood Detection Using Smart Water Sensors
Sub-fields: Hydrology, remote sensing, IoT data science, geospatial analysis, environmental data science, real-time data processing, disaster risk science
Difficulty Level: Intermediate
Research Question: How accurately can smart water sensor data predict flooding in urban neighborhoods?
Many urban areas struggle with flash flooding because drainage systems cannot handle sudden storms, making early warning systems increasingly important. In this project, you can study how connected water sensors and rainfall monitors help cities detect floods before major damage occurs. To conduct your research, you can work with sensor datasets such as river water-level sensors or rainfall monitoring systems, weather records, and real-time monitoring tools such as Arduino-based water sensors, Raspberry Pi systems and online dashboards. You can also use Python libraries such as Pandas or Plotly for trend analysis and data visualization. This topic is a good fit if you are interested in engineering, environmental technology, automation, or smart city systems.
Business Intelligence and Consumer Analytics
This field studies customer behavior, purchasing trends, and product feedback to help businesses improve decisions and user experiences. As a student researcher, you’ll analyze customer reviews, browsing patterns or sales data to identify trends and market opportunities. This field is a great option if you are interested in marketing, entrepreneurship, or technology-driven business research.
29. Studying Consumer Reviews to Improve Product Design
Sub-fields: Natural language processing, sentiment analysis, human-computer interaction, marketing analytics, user experience research
Difficulty Level: Beginner
Research Question: What complaints appear most frequently in smartphone customer reviews?
For this topic, you can analyze product reviews to identify patterns in customer satisfaction and product issues. Companies often rely on review analysis to improve their designs and customer experiences. To conduct your research, you can use review datasets from sources such as Amazon product reviews, Yelp restaurant reviews, or Google Reviews, classify common themes, and visualize trends in customer feedback using Python or spreadsheets. This topic is a good fit if you are interested in business, technology, or beginner natural language processing.
30. Predicting Customer Purchasing Behavior in Online Stores
Sub-fields: Behavioral data science, marketing analytics, recommender systems, consumer psychology, machine learning
Difficulty Level: Beginner
Research Question: Can browsing history predict whether a customer will complete a purchase?
In this project, you can study how online retailers use user behavior data to recommend products and increase sales. Shopping platforms collect large amounts of data, yet predicting purchases remains challenging because customer behavior changes frequently. You can use e-commerce datasets such as Kaggle online retail datasets or Amazon sales records, spreadsheets including Excel or Google Sheets, and simple machine learning models in Python to understand purchasing patterns and investigate how user behavior can be used to predict future purchases. This topic works well if you’re interested in business, marketing, or beginner predictive analytics.
Frequently Asked Questions
Do I need advanced coding skills to complete one of these projects?
No. Several beginner topics, including analyzing income inequality through Census data and predicting movie box office success from online reviews, can be completed using spreadsheets or basic regression analysis in Excel, without advanced programming.
Which topics are best suited for students interested in current events or politics?
Analyzing social media sentiment during elections, predicting voter turnout using demographic data, and measuring the spread of misinformation during natural disasters all connect data science directly to political and civic engagement questions.
What datasets are most commonly used across these topics?
Common sources include Kaggle datasets, the EPA and NOAA for environmental data, the CDC for public health data, the Census Bureau for demographic data, and platform-specific sources like Spotify Charts, IMDb, and Yahoo Finance depending on the topic’s subject area.
How do I know if a topic is too broad or too narrow?
A topic like “social media and mental health” is too broad because it doesn’t specify a measurable question or dataset, while a version like “Analyzing whether Instagram screen time affects sleep patterns among teenagers” is specific enough to define a clear research question, dataset, and methodology.
What if I am interested in research outside of data science?
These topics are focused specifically on data driven research methods, but students looking for a research based credential in a different subject area might consider Horizon’s Research Seminars and Labs, a selective virtual research program that pairs students with a mentor to develop an original research paper across subjects like political theory, biology, and machine learning.
Image source: Horizon Academic Research Program




