Students' academic success in science, technology, engineering, and mathematics (STEM) careers is one of the most popular subjects that has gained attention among educational researchers for decades. Many studies have shown students' educational outcomes can be affected by academic factors including high school GPA, SAT score test, student admission type (transfer or first-time-in-college), as well as demographic features such as gender, ethnicity, and family income. Additional studies have investigated the relationship between students' course load and their academic outcomes. In this paper, we define students' course load based on the number of courses they take each semester, which is assumed to have a discrete probability distribution. To assess if students' course load impacts their academic performance, we apply Hidden Markov models, an unsupervised learning method, to classify students into three categories: high-level enrollment, medium-level enrollment, and low-level enrollment. The sequence of the number of courses students enroll each semester during their academic career is fed into our proposed model as input. The output which is a qualitative measure and is not directly observable is the estimated enrollment level for the students. After students' classification, we derive and compare their academic (e.g., cumulative GPA, graduation rate, and DFW rate) and non-academic (e.g., family income level) features for each enrollment level category. Findings show that students who have more engagement with the university have higher academic performance (higher cumulative GPA and graduation, lower DFW rate) than those with lower engagement. Our analysis also demonstrates that students from families with low-income levels are more likely to have lower enrollment levels. Such results indicate that university managers can improve students' educational performance and, subsequently, the university graduation rate by encouraging students to engage more with the university by providing academic and financial support. These results are based on the data collected from the University of XXX from 2008 to 2016 and contain approximately 170,000 students.
Our study aims to collect data to understand ideological and extreme bias in text articles shared across various online communities, particularly focusing on the language used in subreddits associated with extremism and targeted violence.Initially, we gathered data from related online communities, specifically the r/Liberal and r/Conservative communities on Reddit, utilizing the Reddit Pushshift API to collect URLs shared within these subreddits. Our aim was to gather news, opinion, and feature articles, resulting in a corpus of 226,010 articles. We also curated a balanced subset of 45,108 articles and annotated 4,000 articles to validate their relevance, facilitating understanding of language usage within ideological Reddit communities and insights into ideological bias in media content.Expanding beyond binary ideologies, we introduced a new category termed "Restricted" to encompass articles shared in private or banned subreddits. This third category encompasses articles shared in restricted, privatized, quarantined, or banned subreddits characterized by radicalized and extremist ideologies. This expansion yielded a large dataset of 377,144 articles. Additionally, we included articles from subreddits with unspecified ideologies, creating a holdout set of 922,522 articles. In total, our combined dataset of 1.3 million articles collected from 55 different subreddits will assist in examining radicalized communities and providing discourse analysis in associated subreddits, enhancing understanding of the language used in articles shared within radicalized Reddit communities and offering insights into extreme bias in media content.In summary, we collected 1.52 million articles to understand ideological and extreme bias, providing a comprehensive dataset that aids in understanding language usage within text articles posted in ideological and extreme Reddit communities.
The maturation of autonomy for electric vertical take-off and landing aircraft will soon make it possible to execute military intelligence, surveillance, reconnaissance (ISR) missions aboard crewed autonomous aerial vehicles. This research experimentally investigates factors that may influence the quality of interaction (i.e., team fluency) between a non-pilot human operator and the AI pilot responsible for autonomous flight, aboard a minimally crewed aircraft. In a flight simulator study with twenty-seven participants, various levels of workload and AI pilot capabilities are investigated including run time assurance through control barrier functions (CBFs). CBFs are used to enable pro-active collision avoidance behaviors by the AI pilot. Team fluency and mission effectiveness outcomes through trust, situation awareness, workload, interaction and performance show that task complexity and AI behavior are significant factors for the quality of human AI interaction in the autonomous ISR context.
The main objective of our research is to gain a comprehensive understanding of the relationship between language usage within different communities and delineating the ideological narratives. We focus specifically on utilizing Natural Language Processing techniques to identify underlying narratives in the coded or suggestive language employed by non-normative communities associated with targeted violence. Earlier studies addressed the detection of ideological affiliation through surveys, user studies, and a limited number based on the content of text articles, which still require label curation. Previous work addressed label curation by using ideological subreddits (r/Liberal and r/Conservative for Liberal and Conservative classes) to label the articles shared on those subreddits according to their prescribed ideologies, albeit with a limited dataset.Building upon previous work, we use subreddit ideologies to categorize shared articles. In addition to the conservative and liberal classes, we introduce a new category called “Restricted” which encompasses text articles shared in subreddits that are restricted, privatized, or banned, such as r/TheDonald. The “Restricted” class encompasses posts tied to violence, regardless of conservative or liberal affiliations. Additionally, we augment our dataset with text articles from self-identified subreddits like r/progressive and r/askaconservative for the liberal and conservative classes, respectively. This results in an expanded dataset of 377,144 text articles, consisting of 72,488 liberal, 79,573 conservative, and 225,083 restricted class articles. Our goal is to analyze language variances in different ideological communities, investigate keyword relevance in labeling article orientations, especially in unseen cases (922,522 text articles), and delve into radicalized communities, conducting thorough analysis and interpretation of the results.
Planning and execution of a UAV operation requires properly monitoring and forecasting available power over the complete flight plan. This is especially true in the case of delivery operations where payloads increase power requirements as a battery discharges. In this paper, we document the development and application of a long short-term memory (LSTM) recurrent neural network to predict UAV battery state-of-charge levels as a function the UAV’s flight trajectory. The model is developed and evaluated using over 134,000 achieved real-world UAV flights. Beyond trajectory information, meta-data regarding the aircraft model type, prior discharges, and battery health is integrated into the LSTM model. Through validation exercises we demonstrate that the resulting LSTM model is able to significantly outperform linear models in predicting battery state-of-charge.
Air traffic controllers engage in complex and dynamic decision-making when managing an airspace. This is especially true in the immediate vicinity of an airport. Unlike in en-route or terminal area airspace where aircraft usually traverse well established routes and procedures, near the airport after completing a standard arrival procedure, the routes to the final approach are only partially defined. In this airspace (i.e., 10-12 nautical miles from the airport), the local tower controllers tactically guide aircraft through tromboning and vectoring commands to maintain separation requirements between aircraft and space them out at the runway. In this paper, a Mixed-Integer Linear Programming formulation is used to design the order-sequencing of aircraft at a single runway, by allowing vectoring and tromboning, to maximize airport throughput and minimize the distance traversed while maintaining safety. The mathematical model is formulated by time-metering aircraft at potential conflict points to avoid conflicts. With the goal of building a decision support tool that emulates the control techniques and tactical maneuvers of local tower controllers, the proposed optimization formulation generates conflict-free and safe trajectories that conform to the scheduling.
The paper proposes a generative pedestrian trajectory modeling framework named HISS - Human Interactions in Shared Space. The trajectory modeling framework is based on a receding horizon optimization approach utilizing pedestrian behavior and interactions that seeks to capture pedestrian trajectory planning and execution. The benefit of the proposed dynamic optimization trajectory generation approach is that it requires minimal calibration data under a variety of traffic scenarios. In this paper, we formalize several pedestrian-pedestrian interaction scenarios and implement trajectories’ conflict avoidance through mixed integer linear programming (MILP). We validate the proposed framework on two benchmark datasets - DUT and TrajNet++. The paper shows that when the framework’s parameters are tuned to certain initial conditions and pedestrian behavior and interaction rules, the framework generates pedestrian trajectories similar to those observable in real-world scenarios, justifying the framework’s capability to provide explanations and solutions to various traffic situations. This feature makes the proposed framework useful for modelers and urban city planners in making policy decisions.
This research addresses the crucial challenge of effectively measuring threats in social media comments targeting voting, public officials, and institutions in the United States. Our understanding of these online threats and their links to real-world risks is limited, making it difficult to assess their seriousness. To overcome these limitations, we propose a comprehensive threat level scale from 0 to 5 and collect a dataset of 1.3 million Telegram responses for developing and rigorously testing these threat levels. Additionally, we explore OpenAI-human annotation to efficiently label this vast dataset. Our innovative two-step transfer learning approach initially employs a pre-existing, pre-trained model for labeling, followed by expert validation. Next, we use the AI-annotated samples to develop independent models, and expert annotators verify their predictions. Notably, our findings demonstrate that the GPT-2 model, despite its fewer annotated training set, performs comparably to OpenAI's anno-tations, showcasing its potential for cost-effective threat detection with more annotated samples. With the long-term objective of establishing continuous threat-level monitoring, we identify the strengths and limitations of our current approach and propose a roadmap for enhancing threat detection.
View Video Presentation: https://doi.org/10.2514/6.2022-3709.vid In this paper, the authors seek to identify potential inefficiencies at airports controlled by local tower controllers when sequencing aircraft. Using naturalistic data in the form of historical air traffic radar data, we create and evaluate statistical and machine learning models of air traffic controllers when ordering aircraft to land at the airport. The models are based on binary decision classifiers (e.g., logistic regression) that take as their input static information related to the entrance times of aircraft, their traffic flow pattern, and dynamic state information (position, speed, etc) in order to predict a relative ordering. Once the relative ordering between all aircraft is established, a landing sequence is induced. A comparison of model predictions against the actualized decisions by air traffic controllers indicates that it is unlikely that air traffic controllers are making use of additional state information beyond an initial entry time and estimated landing time for each aircraft when setting a landing sequence. Furthermore, when performing the comparison on a set of historical trajectories at Ronald Reagan Washington National Airport (DCA) we demonstrate that mismatches between actualized decisions and the prediction models is associated with excess distance traversed by the aircraft, as such it appears there is room for additional sequence optimization.
With the long-term goal of understanding how language is used and evolves within online communities, this work explores the application of natural language processing techniques to classify text articles according to their ideological orientation (i.e., conservative or liberal). We first collect a balanced corpus of text articles posted to the online communities r/Liberal and r/Conservative from the social media website Reddit. Using the corpus, we develop and apply three classifiers. The baseline classifier is a Bayes model that accounts for each text article’s web domain, as such, classification is independent of content. Next, we develop a support vector machine (SVM) model with term frequency-inverse document frequency (TF-IDF) features; this approach highlight differences in language using a count-based feature-space to differentiate text articles. Last, we evaluate the context-based transformer (RoBERTa) model and discuss its under-performance relative to the baseline and SVM models.
Many studies in the field of education analytics have identified student grade point averages (GPA) as an important indicator and predictor of students' final academic outcomes (graduate or halt). And while semester-to-semester fluctuations in GPA are considered normal, significant changes in academic performance may warrant more thorough investigation and consideration, particularly with regards to final academic outcomes. However, such an approach is challenging due to the difficulties of representing complex academic trajectories over an academic career. In this study, we apply a Hidden Markov Model (HMM) to provide a standard and intuitive classification over students' academic-performance levels, which leads to a compact representation of academic-performance trajectories. Next, we explore the relationship between different academic-performance trajectories and their correspondence to final academic success. Based on student transcript data from University of Central Florida, our proposed HMM is trained using sequences of students' course grades for each semester. Through the HMM, our analysis follows the expected finding that higher academic performance levels correlate with lower halt rates. However, in this paper, we identify that there exist many scenarios in which both improving or worsening academic-performance trajectories actually correlate to higher graduation rates. This counter-intuitive finding is made possible through the proposed and developed HMM model.
Urbanization is bringing together various modes of transport, and with that, there are challenges to maintaining the safety of all road users, especially vulnerable road users (VRUs). There is a need for street designs that encourages cooperation between road users. Shared space is a street design approach that softens the demarcation of vehicles and pedestrian traffic by reducing traffic rules, traffic signals, road marking, and regulations. Understanding the interactions and trajectory formations of various VRUs will facilitate the design of safer shared spaces. In line with this goal, this paper aims to develop a methodology for generating VRUs trajectories that accounts for behaviors and social interactions. We develop a receding horizon optimization-based pedestrian trajectory planning algorithm capable of modeling pedestrian trajectories in a variety of shared space scenarios. Focusing on three scenarios-group interactions, unidirectional interaction, and fixed obstacle interaction-case studies are performed to demonstrate the strengths of the resulting generative model. Additionally, generated trajectories are validated using two benchmark datasets – DUT and TrajNet++. The three case studies are shown to yield low or near-zero Mean Euclidean Distance and Final Displacement Error values supporting the performance validity of the models. We also analyze gait parameters (step length and step frequency) to further demonstrate the model’s capability at generating realistic pedestrian trajectories.
Simplified classifications have often led to college students being labeled as full-time or part-time students. However, student enrollment patterns can be much more complicated at many universities, as it is common for students to switch between full-time and part-time enrollment each semester based on finances, scheduling, or family needs. While previous studies have identified part-time enrollment as a risk factor to students’ academic success, limited research has examined the impact of enrollment patterns or strategies on academic performance. Unlike traditional methods that use a single-period model to classify students into full-time and part-time categories, in this study, we apply an advanced multi-period dynamic approach using a Hidden Markov Model to distinguish and cluster students’ enrollment strategies into three categories: full-time, part-time, and mixed. We then investigate and compare the academic performance outcomes of each group based on their enrollment strategies while taking into account student type (i.e., first-time-in-college students and transfer students). Analysis of undergraduate student records data collected at the University of Central Florida from 2008 to 2017 shows that the academic performance of first-time-in-college students who apply a mixed enrollment strategy is closer to that of full-time students, as compared to part-time students. Moreover, during their part-time semesters, mixed-enrollment students significantly outperform part-time students. Similarly, analysis of transfer students shows that a mixed-enrollment strategy is correlated with similar graduation rates as the full-time enrollment strategy and more than double the graduation rate associated with part-time enrollment. This finding suggests that part-time students can achieve better overall outcomes by increased engagement through occasional full-time enrollments.
In an effort to maintain safety while satisfying growing air traffic demand, air navigation service providers are considering the inclusion of advisory systems to identify potential conflicts and propose resolution commands for the air traffic controller to verify and issue to aircraft. To understand the potential workload implications of introducing advisory conflict-detection and resolution tools, this paper examines a metric of controller taskload: how many resolution commands an air traffic controller issues under the guidance of an advisory system. Through a simulation study, the research presented here evaluates how the underlying protocol of a conflict-resolution tool affects the controller taskload (system demands) associated with the conflict-resolution process, and implicitly the controller workload (physical and psychological demands). Ultimately, evidence indicates that there is significant flexibility in the design of conflict-resolution algorithms supporting an advisory system.
Recent advancements in machine learning and the availability of massive data sets now facilitate the development of neural network models to predict aviation noise based on aircraft trajectories. In this paper, we use a long short-term memory recurrent neural network to predict aviation noise at a ground location near Washington National Airport. The model is developed using over 10 months of achieved radar data and noise readings. Beyond trajectory information, meta-data regarding the aircraft type and weather data is integrated into the model. While specific to a ground station and airport, the model is able to accurately predict aviation noise above 55 db with a mean absolute error of 2.3 dB.
Accurate trajectory prediction is required to realize safe and efficient aircraft operations. In this paper, a new framework for predicting arrival time of en-route aircraft using Gaussian Mixture Model (GMM) is proposed. The proposed method fits the historical trajectory data with GMM whose variable is a set of arrival times at the significant points along a specific air route. The flight times to the defined points along the air route are computed conditioned on the observed flight times for the previous points that the aircraft has already passed by. The form of prediction output from the proposed model is the probability distribution which would increase its applicability to various fields due to its probabilistic nature. The performance of the proposed method is demonstrated by applying it to real flight data in Incheon Flight Information Region (FIR). Keywords-component; Trajectory Prediction; Departure Manager; Gaussian Mixture Model; Probabilistic Model; Unsupervised Learning
College students are enrolled at each semester with either part time or full time status. While most of the students keep an overall constant enrollment status during their education period, some of them may frequently change their status between full time and part time from one semester to the next. The goal of this research is to exploit the historic patterns to estimate and categorize students$'$ strategy in three different groups of part time, full time and mixed, investigate the educational features of each group and compare their performance. Enrollment strategy refers to the student$'$s mindset for enrollment plan and in one way can be captured from the student$'$s historic enrollment status. Data is collected from the University of Central Florida from 2008 to 2017 and Hidden Markov Model is applied to identify different types of student strategy. Results show that students with Mixed Enrollment Strategy (MES) have features (ex. time to graduation and graduation and halt enrollment ratio) and performances (ex. cumulative GPA) relatively between students with Full time Enrollment Strategy (FES) and students with Part time Enrollment Strategy (PES).
Simplified categorizations have often led to college students being labeled as full-time or part-time students. However, at many universities student enrollment patterns can be much more complicated, as it is not uncommon for students to alternate between full-time and part-time enrollment each semester based on finances, scheduling, or family needs. While prior research has established full-time students maintain better outcomes then their part-time counterparts, limited study has examined the impact of enrollment patterns or strategies on academic outcomes. In this paper, we applying a Hidden Markov Model to identify and cluster students' enrollment strategies into three different categorizes: full-time, part-time, and mixed-enrollment strategies. Based the enrollment strategies we investigate and compare the academic performance outcomes of each group, taking into account differences between first-time-in-college students and transfer students. Analysis of data collected from the University of Central Florida from 2008 to 2017 indicates that first-time-in-college students that apply a mixed enrollment strategy are closer in performance to full-time students, as compared to part-time students. More importantly, during their part-time semesters, mixed-enrollment students significantly outperform part-time students. Similarly, analysis of transfer students shows that a mixed-enrollment strategy is correlated a similar graduation rates as the full-time enrollment strategy, and more than double the graduation rate associated with part-time enrollment. Such a finding suggests that increased engagement through the occasional full-time enrollment leads to better overall outcomes.
American universities use a procedure based on a rolling six-year graduation rate to calculate statistics regarding their students’ final educational outcomes (graduating or not graduating). As an alternative to the six-year graduation rate method, many studies have applied absorbing Markov chains for estimating graduation rates. In both cases, a frequentist approach is used. For the standard six-year graduation rate method, the frequentist approach corresponds to counting the number of students who finished their program within six years and dividing by the number of students who entered that year. In the case of absorbing Markov chains, the frequentist approach is used to compute the underlying transition matrix, which is then used to estimate the graduation rate. In this paper, we apply a sensitivity analysis to compare the performance of the standard six-year graduation rate method with that of absorbing Markov chains. Through the analysis, we highlight significant limitations with regards to the estimation accuracy of both approaches when applied to small sample sizes or cohorts at a university. Additionally, we note that the Absorbing Markov chain method introduces a significant bias, which leads to an underestimation of the true graduation rate. To overcome both these challenges, we propose and evaluate the use of a regularly updating multi-level absorbing Markov chain (RUML-AMC) in which the transition matrix is updated year to year. We empirically demonstrate that the proposed RUML-AMC approach nearly eliminates estimation bias while reducing the estimation variation by more than 40%, especially for populations with small sample sizes.
The inclusion of large-scale data-monitoring and data-storage systems has presented researchers with new opportunities for understanding the decisions of air traffic controllers in managing traffic. This paper takes a step in the direction of data-driven analysis by clustering trajectories according to air traffic control decisions (e.g. trajectory changes), as opposed to clustering on spatial positioning reports. The particular focus for this paper is in the immediate vicinity of airports. Unlike in enroute or terminal area airspace where aircraft typically traverse well established routes and procedures, near airports there tends to be greater variability in air traffic decisions (e.g. vectoring and tromboning) to appropriately space aircraft at the runway, ultimately leading to spatial dispersion of aircraft trajectories. Using a hidden Markov model, heading changes within aircraft trajectories are identified, extracted, and translated into variable-length trajectory strings. Comparison of the trajectory strings using edit distance metrics then allows for clustering of the trajectories according to the air traffic control decisions. The classification and clustering process is applied on a set of historical trajectories at Washington National Airport. The resulting clusters foster an understanding of the arrival traffic structure and the decision strategies of controllers.