This paper presents a case study of how to use machinelearning to provide actionable recommendations in themilitary assessment and selection (A&S) context. Militaryunits use an in-depth A&S process to acquire the mostqualified candidates. The specific A&S we studied had11,885 candidate records collected over afive-year periodwith 89 total features that included administrative, per-formance, and psychological data on each candidate. Weapplied a robust machine learning approach involvingfeature engineering, feature selection, optimized predic-tive models, and data subsets analysis to extract meaning-ful information from the data. Our objective for thisresearch was to evaluate the utility of applying machinelearning techniques to a specific military A&S datasetwith the goal of improving the holistic A&S selection pro-cess. The research resulted in valuable insights from thedata that found features highly predictive of candidatenonselection, learned methods to modify the existing datato improve predictive capability, and informed actionablerecommendations for the A&S selection process.
We propose a novel asset allocation model using a Markov process of states defined by clustered efficient frontier coefficients. While most research in Markov models of the market characterize regimes using return and volatility, we instead propose characterizing these states using efficient frontiers, which provide more information on the interactions of underlying assets that comprise the market. Efficient frontiers can be decomposed to their functional form, a square-root second-order polynomial defined by three coefficients, to provide a dimensionality reduction of the return vector and covariance matrix. Each month, the proposed model hierarchically clusters the monthly coefficients data up to the current month, to characterize the market states, then defines a Markov process on the sequence of states. To incorporate these states into portfolio optimization, for each state, we calculate the tangency portfolio using only return data in that state. We then take the expectation of these weights for each state, weighted by the probability of transitioning from the current state to each state. To empirically validate our proposed model, we employ three sets of assets that span the market, and show that our proposed model significantly outperforms benchmark portfolios.
The position paper summarizes the inputs of a set of experts from academia and industry presenting their view on chances and challenges of using ChatGPT within the Modelling and Simulation education. They also address the need to evaluate continuous education as well as education of faculty members to address scholastic challenges and opportunities while meeting the expectation of industry.
Since the first use of the term "systems engineering" in the 1940’s, the discipline has progressed significantly as the complexity of systems and the development of technology continue to increase. This paper examines the current state and evolution of systems engineering and systems engineering education. The paper first identifies the current state of systems engineering by recognizing systems engineers' roles, responsibilities, and expectations. Following the current state of systems engineering, the changes needed are addressed through multiple frameworks and skill sets that will be critical to the evolving industry. The paper also explores the future of systems education focusing on content, content delivery, cost, and student cooperation. The analysis suggests that universities could adjust their curriculum to better align with the demands of the industry. The paper concludes with an overview of potential solutions designed to meet the needs of systems engineers by preparing them for the growing and multifaceted industry demands of the future.
We propose a novel model to achieve superior out-of-sample Sharpe ratios. While most research in asset allocation focuses on estimating the return vector and covariance matrix, the first component of our novel model instead forecasts the future tangency portfolio, and the second component then determines the optimal investment portfolio. First, to forecast the tangency portfolio, we forecast the efficient frontier by decomposing its functional form, a square root second-order polynomial, into three interpretable coefficients, which can then be used to calculate a forecasted tangency portfolio. These coefficients can be forecasted using vector autoregressions. Second, the model invests in the portfolio on the efficient frontier that is the minimum Euclidean distance from this forecasted tangency portfolio. A motivation for our approach is to address the limitation that the tangency portfolio only maximizes the Sharpe ratio when future returns and covariances are stationary, and can be directly estimated with historical data, which often does not hold in out-of-sample data. Our approach addresses this shortcoming in a novel way by forecasting the tangency portfolio, rather than estimating return and covariance. For empirical testing, we employ two sets of assets that span the market to demonstrate and validate the performance of this novel method.
Recent efforts have shown that training data is not secured through the generalization and abstraction of algorithms. This vulnerability to the training data has been expressed through membership inference attacks that seek to discover the use of specific records within the training dataset of a model. Additionally, disparate membership inference attacks have been shown to achieve better accuracy compared with their macro attack counterparts. These disparate membership inference attacks use a pragmatic approach to attack individual, more vulnerable sub-sets of the data, such as underrepresented classes. While previous work in this field has explored model vulnerability to these attacks, this effort explores the vulnerability of datasets themselves to disparate membership inference attacks. This is accomplished through the development of a vulnerability-classification model that classifies datasets as vulnerable or secure to these attacks. To develop this model, a vulnerability-classification dataset is developed from over 100 datasets—including frequently cited datasets within the field. These datasets are described using a feature set of over 100 features and assigned labels developed from a combination of various modeling and attack strategies. By averaging the attack accuracy over 13 different modeling and attack strategies, the authors explore the vulnerabilities of the datasets themselves as opposed to a particular modeling or attack effort. The in-class observational distance, width ratio, and the proportion of discrete features are found to dominate the attributes defining dataset vulnerability to disparate membership inference attacks. These features are explored in deeper detail and used to develop exploratory methods for hardening these class-based sub-datasets against attacks showing preliminary mitigation success with combinations of feature reduction and class-balancing strategies.
Social Network Services (SNS) are systems that allow users to build social relations with one another, with one of the largest SNSs being Facebook, totaling 2.6 billion active monthly users in 2020. Their live streaming service, Facebook Live, is one of the fastest-growing branches of the company, allowing creators to synchronously broadcast original content to the public. However, in the rapidly growing world of technology, Facebook Live faces fierce competition from other live streaming platforms (Twitch, YouTube Live, etc.) and well as other video-on-demand providers (Netflix, Hulu, TikTok, etc.). To better understand current issues and future directions, our team focused on the Facebook Live platform, to develop a three to five-year strategic plan for the platform. We focus on Facebook Live's growth opportunities from a multitude of perspectives, including the competitive landscape, interface modification, future projections based on historical trends, and competitive analysis. Our approach utilizes the systems analysis process, focusing top-down on objectives and metrics. To produce a comprehensive strategy for future operations, we employ analytical methods ranging from quantitative data analysis to qualitative exploration of industry trends. These quantitative methods include statistical analysis, time-series forecasting, and natural language processing. Qualitative methods include domain research into the history, current state, and possible future for live-streaming. Forthcoming, the results for the complete analysis will be synthesized into a multi-recommendation strategic report to provide Facebook with flexible guidance for continuing operations. We also present a comments summarization and visualization feature for viewers and creators, three to five-year market forecasts after COVID-19 lockdowns, and attractive emerging markets including education and morning shows.
The Markowitz model is an established approach to portfolio optimization that constructs efficient frontiers allowing users to make optimal tradeoffs between risk and return. However, a limitation of this approach is that it assumes future asset returns and covariances will be identical to the asset's historical data, or that these model parameters can be accurately estimated, a notion which often does not hold in practice. Markowitz efficient frontiers are square root second-order polynomials that can be represented by three parameters, thus providing a significant dimensionality reduction of the lookback covariances and growth of the assets. Using this dimensionality reduction, we propose an extension to the Markowitz model that accounts for the nonstationary behavior of the portfolio assets' return and covariance without the necessity to forecast the complex covariance matrix and assets growths, something that has proven to be extremely difficult. Our methodology allows users to forecast the three efficient frontier coefficients using a time-series regression. By observing similar efficient frontiers, this forecasted efficient frontier can be used to select optimal assets mean-variance tradeoffs (asset weights). For exploratory testing we employ a set of assets that span a large portion of the market to demonstrate and validate this new approach.
Organizations in the nonprofit space are increasingly using data mining techniques to gain insights into their donors’ behaviors and motivations. Data mining can be costly but can also be valuable in retaining and obtaining donors. Throughout the course of this project, we have prioritized two objectives. One is to increase the ratio of funds raised to dollars spent on fundraising from current donors, making these efforts more profitable. The other is to determine how to most effectively solicit new donors. To accomplish these goals, we have used statistical modeling and data analysis to gain insights and create recommendations related to donor optimization and acquisition. To learn about the current donors, it is important to identify which unique traits make donors more likely to donate and whether those traits are related to an individual’s demographic information or giving history. Our team is classifying donors into "states" of giving based upon different metrics, including how recently, how much, how often, and for how long they have donated. We are using various data models to create actionable recommendations on how to tailor fundraising appeals specifically to different donors, which will increase the Inn’s overall donations and their return on fundraising investment. We are also mapping the transitions between these giving states so that donors dropping from higher states can be re-engaged, while donors with a high chance of moving into a more profitable state can be flagged and targeted. We will present these results in a dashboard that the Inn can use moving forward to better solicit each donor and maintain a steady fundraising revenue stream.
In the world of college sports, the process of recruiting players is one of the most important tasks a coach must tackle. With only 6% of the 8 million high school athletes earning spots on NCAA teams, finding and selecting the right players can be incredibly challenging even with the availability of widespread data. Some sports, like football and basketball, have found great success using predictive analytics to estimate success in college. These efforts, however, have not yet been extended to other sports, such as golf. Given the vast amount of data available to the public on junior golfers, there is clear potential to bring analytics to college golf recruiting. We partnered with GameForge, a leading golf analytics company, to create a recommendation tool for college coaches, one that leverages the already existing data on high school and collegiate golfers and a variety of predictive models to display athletes we believe would best fit in a certain college program. A systems analysis approach was taken to find the factors that most accurately predict a high school player's success in college golf. This was done with a variety of models including the forecasting of probability of a high school athlete being a top ranked college golfer, the finding of players with a similar performance to another desired player, and the predicting of a junior golfer's scoring performance and development during the remainder of their high school career and during college. Using these models, we identified several factors that are predictive of player similarity and performance. The research team iteratively developed these models to be used in conjunction with each other in order to provide meaningful, and understandable recommendations to a college coach on which players they should recruit to maximize success.
Cognitive state detection and its relationship to observable physiologically telemetry has been utilized for many human-machine and human-cybernetic applications. This paper aims at understanding and addressing if there are unique psychophysiological patterns over time, a ''physiological temporal fingerprint'', that is associated with specific cognitive states. This preliminary work involves commercial airline pilots completing experimental benchmark task inductions of three cognitive states: 1) Channelized Attention (CA); 2) High Workload (HW); and 3) Low Workload (LW). We approach this objective by modeling these "fingerprints" through the use of Hidden Markov Models and Entropy analysis to evaluate if the transitions over time are complex or rhythmic/predictable by nature. Our results indicate that cognitive states do have unique complexity of physiological sequences that are statistically different from other cognitive states. More specifically, CA has a significantly higher temporal psychophysiological complexity than HW and LW in EEG and ECG telemetry signals. With regards to respiration telemetry, CA has a lower temporal psychophysiological complexity than HW and LW. Through our preliminary work, addressing this unique underpinning can inform whether these underlying dynamics can be utilized to understand how humans transition between cognitive states and for improved detection of cognitive states.
We discuss the program-level model used in the administration of undergraduate Capstone (senior design) projects in the Department of Systems and Information Engineering at University of Virginia’s School of Engineering and Applied Science in this paper. A unique model at the time of its inception in 1988, its adoption by other institutions and its longevity are measures of its effectiveness and robustness. We provide an overview of the tasks performed by the various personnel involved in the administration of the undergraduate Capstone projects in chronological order, starting with activities performed during the summer recess, after a brief introduction to our Department’s Capstone Program. The goal of providing this “cookbook” review of the model is to provide sufficient information to allow other departments to adopt similar practices.
This paper presents an innovative new approach to investment portfolio design, which applies a discrete, state-based methodology to defining market states and making asset allocation decisions with respect to both current and future state membership. State membership is based on attributes taken from traditional finance and portfolio theory namely expected growth, and covariance. The transitional dynamics of the derived states are modeled as a Markovian process. Asset weighting and portfolio allocation decisions are made through an optimization-based approach coupled with heuristics that account for the probability of state membership and the quality of the state in terms of information provided.
at the University of Virginia, Charlottesville, Virginia.CSER 2018 offered researchers in academia, industry, and government a common forum to present, discuss, and influence systems engineering research.The conference provided access to forward-looking research by renowned academicians as well as perspectives from senior industry and government representatives.Co-founded by the University of Southern California and Stevens Institute of Technology in 2003, the CSER series of conferences has become the preeminent event for researchers in systems engineering across the globe.The conference theme of Systems Engineering in Context was intended to highlight the ways that context-economic, social, cultural, organizational, and otherwise-can shape decision problems and other aspects of systems engineering.In their introduction to the inaugural issue of Environment Systems and Decisions, Linkov, Lambert, and Collier (2013) stated the first word in the title of the journal "is intended to imply not only the natural environment, but also the environment in which a problem, decision, or innovation exists.
As we delve further into the data science/big data era, we create more and more forecasts or predictions for systems evaluation applications. Many practitioners use ratio variables as performance metrics, one of the most common being the benefit/cost (B/C) ratio of alternatives; however, ratio metrics are used extensively throughout disparate industries. There is significant risk associated with using ratio metrics, and in this paper, we focus on one simple trap that can result from forecasting a ratio metric. We illustrate and motivate the issues with a stylized television-viewing index forecasting problem and with two examples using real-world data. Our goal in this paper is to make systems practitioners aware of these pitfalls.
Individualized biometric data are being incorporated into training and competitions by many coaches and trainers to provide insights into athletic performance and physical fitness of their athletes. Currently, fitness tracking software provides coaches with minimal descriptive statistics on the collected biometric data, resulting in limited actionable outcomes. The collection of biometric data provides an opportunity to understand the variables that are indicative of athletic performance, and to create predictive models to determine appropriate training and in-game strategies. In order to develop these informative decision support tools, predictive frameworks have to address the correct performance metrics, control of subject-to-subject variability, handle data limitations, and maintain model interpretability. We demonstrate that the strenuousness of training sessions leading up to a competitive match has significant impact on the outcome of the game (win or loss) in continuous-play team sports. Specifically, a high cardiovascular training load two days prior to competition was predictive of a win. Additionally, we show that statistically significant differences exist in the physiological behaviors of different player positions. Analysis of several performance metrics also demonstrates that singular metrics or combinations of simple statistics do not directly relate to the outcome of a game, particularly in low-scoring sports such as field hockey or soccer.
Machine learning is a collection of techniques designed to detect hidden patterns in data and construct or learn models for predicting the outcome of future data. These predictions can be used to make decisions about the system generating the data in the face of uncertainty. This chapter discusses a mathematical formulation for incorporating insights, intuition, or prior knowledge into the modeling process using informative prior distributions. Prior knowledge can also influence the modeling process. Machine-learning models can often make assumptions about the process or the data. One area of machine learning is focused on learning parametric models that explain the behavior of the data. These types of parametric models often make assumptions about how the data interact with an outcome or the distributions generating the data. Data-driven modeling has become very popular in the last decade. Neural networks and deep-learning techniques have been shown to be extremely powerful in problems such as supervised classification of images.
The Veterans Health Administration (VHA) is plagued by abnormally high no-show and cancellation rates that reduce the productivity and efficiency of its medical outpatient clinics. We address this issue by developing a dynamic scheduling system that utilizes mobile computing via geo-location data to estimate the likelihood of a patient arriving on time for a scheduled appointment. These likelihoods are used to update the clinic's schedule in real time. When a patient's arrival probability falls below a given threshold, the patient's appointment is canceled. This appointment is immediately reassigned to another patient drawn from a pool of patients who are actively seeking an appointment. The replacement patients are prioritized using their arrival probability. Real-world data were not available for this study, so synthetic patient data were generated to test the feasibility of the design. The method for predicting the arrival probability was verified on a real set of taxicab data. This study demonstrates that dynamic scheduling using geo-location data can reduce the number of unused appointments with minimal risk of double booking resulting from incorrect predictions. We acknowledge that there could be privacy concerns with regards to government possession of one's location and offer strategies for alleviating these concerns in our conclusion.
Using an agent-based model of the limit order book, we explore how the levels of information available to participants, exchanges, and regulators can be used to improve our understanding of the stability and resiliency of a market. Ultimately, we want to know if electronic market data contains previously undetected information that could allow us to better assess market stability. Using data produced in the controlled environment of an agent-based model’s limit order book, we examine various resiliency indicators to determine their predictive capabilities. Most of the types of data created have traditionally been available either publicly or on a restricted basis to regulators and exchanges, but other types have never been collected. We confirmed our findings using actual order flow data with user identifications included from the CME (Chicago Mercantile Exchange) and New York Mercantile Exchange. Our findings strongly suggest that high-fidelity microstructure data in combination with price data can be used to define stability indicators capable of reliably signaling a high likelihood for an imminent flash crash event about one minute before it occurs.