We propose a novel method to improve estimation of asset returns for portfolio optimization. This approach first performs a monthly directional market forecast using an online decision tree. The decision tree is trained on a novel set of features engineered from portfolio theory: the efficient frontier functional coefficients. Efficient frontiers can be decomposed to their functional form, a square-root second-order polynomial, and the coefficients of this function captures the information of all the constituents that compose the market in the current time period. To make these forecasts actionable, these directional forecasts are integrated to a portfolio optimization framework using expected returns conditional on the market forecast as an estimate for the return vector. This conditional expectation is calculated using the inverse Mills ratio, and the Capital Asset Pricing Model is used to translate the market forecast to individual asset forecasts. This novel method outperforms baseline portfolios, as well as other feature sets including technical indicators and the Fama-French factors. To empirically validate the proposed model, we employ a set of market sector ETFs.
Scientific misconduct has emerged as a growing risk to the academic knowledge base. Questionable research practices such as falsified peer review, predatory conferences, and citation gaming in journal publications have become more prevalent in recent years. As researchers face intense pressure to publish quickly amidst the demand for scholarly findings and literature, the underlying structure of the publishing and research system promotes opportunities for misconduct. The publish-orperish culture creates incentives for scholars, institutions, and journals to engage in questionable behavior, threatening scientific integrity and public welfare. First, this project synthesizes and classifies the scale of scholarly misconduct in academia during the digital age through a comprehensive taxonomy of questionable research practices. Through a literature review and conversations with library science experts, the types of scientific misconduct were classified in a hierarchical taxonomy. This taxonomy was categorized by perpetrator and type of misconduct. The taxonomy and scope were validated through subject matter expert review. An assessment of the scope of each threat was performed using descriptive statistics and time series quantitative data analysis. This analysis was used to identify trends and inform future work.
With the growing integration of artificial intelligence (AI) into autonomous decision-making, ensuring trust in these complex systems is crucial, particularly in life-critical applications where failures can be catastrophic. Existing AI-driven autonomous technology often operates under high uncertainty due to its black-box nature, demanding greater accountability, reliability, and transparency for mission success. This paper proposes a generalizable systems engineering framework for building trust in autonomous systems, demonstrated in the context of minefield traversal, a life-critical control problem. By integrating explainable statistical models into reinforcement learning (RL), this approach evaluates subsystem accuracy and uncertainty in real time, significantly enhancing reliability. Mine detection is supported by two independent, imperfect predictors, an AI model and a human evaluator, each affected differently by varying environmental conditions. Statistical methods quantify prediction reliability, while RL optimizes decisions under uncertainty. Embedding explainable statistics into RL decision-making ensures interpretable outcomes, robust risk-based monitoring, and adaptability to changing operational parameters. This approach was tested through an agent-based simulation where AI and human detection systems collaboratively navigated uncertain minefields. Results indicate improved decision transparency, AI adaptability, and real-time risk management. Explicitly designed for generalizability, this framework presents a scalable method to establish reliable autonomous systems across various safety-critical domains. Future work will refine trust metrics and explore applications in a broader context.
GolfCask, a newly established technology-based start-up in Charlottesville, Virginia, is dedicated to cultivating a vibrant online community centered around the shared passions of golf, travel, and whiskey. Employing a systems analysis framework, this project leverages data-driven insights to refine performance indicators and enhance system efficiency within the online community. The strategic objectives, initially structured within an objective tree, prioritize establishing a sustainable and profitable business model, fostering strong community engagement, and attracting and retaining members. A key component involves designing, validating, and deploying an innovative recommendation system to match user profiles with customized whiskey suggestions. Additional tactics in marketing and data processing are implemented and guided by data analytics and visualization tools to strengthen the technology development and enhance user engagement. Research on user acceptance testing and data integration supplies further insight into developing a user-centric design. Incorporating feedback loops and continuous data analytics refines system outputs, ensuring the recommendations and marketing strategies align with user preferences and business objectives. Integrating a comprehensive system analysis with personalized recommendations and marketing strategies, this project seeks to evaluate the effectiveness of these approaches and provide actionable recommendations for GolfCask’s future technical and business developments.
The STEM field is unrepresentative of the population it serves. Due to a lack of cultural relevance in STEM courses, there is a dissociation between the lived experience of students from underrepresented racial groups (URG) and STEM course material. The SPORT-C intervention is a framework that combines sports, systems thinking learning, and a case-based pedagogy into an activity that can be used in any STEM course. A pilot study was conducted to determine the viability of the SPORT-C intervention in a classroom setting and determine if it was worth further investigating and if any impact differed by racial identity. The findings from this study implicate that the SPORT-C intervention has an impact on the motivation levels of students to participate in STEM courses.
The college sports industry has grown tremendously over the past decade, with NCAA athletic departments recruiting almost half-a-million students to 19,866 teams in 2019 and generating $18.9 billion of revenue the same year. Identifying and selecting the best student-athletes is critical to maintaining the power of these sports programs, aggrandizing the recruitment pipeline and necessitating the demand for novel use of existing technologies. Sports analytics is one response to these growing needs, as its primary use in junior recruitment has presented fruitful for college basketball and football teams across the nation. Golf analytics firm GameForge aims to provide the same insights to college golf coaches, streamlining the recruitment of junior golfers to U.S. universities from around the world. GameForge seeks to develop a two-sided recruiting system that provides insights to junior players and their coaches as well as strengthen its predictive models with the inclusion of new data. A systems-based approach was taken to develop data-driven machine learning models that would provide (a) a proprietary ranking system that compares junior athletes to one another; (b) a relative SWOT analysis that highlights each player's strengths and skill gaps; and (c) a recommender system that suggests potential recruits to college coaches and recommends colleges of best fit to junior players.
Quantifiable, measurable risk is of critical importance when making data-driven decisions in finance and investment management, but what if the generally accepted practice of the investment industry for calculating risk possessed incorrect mathematical assumptions and embedded biases? This piece revisits the discussion surrounding the methodology used to calculate annualized standard deviation statistics commonly used when reporting the performance of investment products. It goes on to present a new example illustrating the bias when applied to an efficient frontier.
The NBA, MLB, NFL and other professional leagues utilize sports analytics, but the potential of professional golf analytics is largely untapped. Instead of using data-driven methods connecting practice to tournament performance, training regimens are often based on conventional wisdom. How can data be used to recommend training regimens for golfers to improve performance? We partnered with golf analytics company, GameForge, to develop tools and methods for golf analytics to capture these markets, including the development of a state-based training recommendation system. We used Gameforge, PGA, and LPGA data to build markov models using k-means clustering, and linear models. These two model types form the basis of our recommendation system. In the future, these methods can be used to inform training decisions, particularly as more data is collected.
As of 2019, sports analytics has grown to be a $780 million industry [1]. Many organizations and institutions contribute to the field through research in exercise science, optimization of in-game decision making, sports marketing, business performance, and sports compliance fields. We propose an open, interdisciplinary approach to sports analytics within institutes of higher education to work across many fields and provide opportunities to diverse members within the community, enable research and communication across fields, serve the surrounding community, and ethically use data.
Systems EngineeringVolume 22, Issue 5 p. 369-369 EDITORIAL Research Advances with Systems Engineering in Context Stephen Conway Adams, Stephen Conway Adams orcid.org/0000-0002-1207-4504 Search for more papers by this authorPeter Beling, Peter BelingSearch for more papers by this authorWilliam Scherer, William SchererSearch for more papers by this authorCody Fleming, Cody FlemingSearch for more papers by this authorJames H. Lambert, James H. Lambert sca2c@virginia.edu Search for more papers by this author Stephen Conway Adams, Stephen Conway Adams orcid.org/0000-0002-1207-4504 Search for more papers by this authorPeter Beling, Peter BelingSearch for more papers by this authorWilliam Scherer, William SchererSearch for more papers by this authorCody Fleming, Cody FlemingSearch for more papers by this authorJames H. Lambert, James H. Lambert sca2c@virginia.edu Search for more papers by this author First published: 10 September 2019 https://doi.org/10.1002/sys.21510Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinkedInRedditWechat No abstract is available for this article. Volume22, Issue5CSER Special IssueSeptember 2019Pages 369-369 RelatedInformation
The role that data analytics plays on sports teams has increased dramatically since Michael Lewis wrote Moneyball and shed some light on Billy Beane's use of analytics with the Oakland Athletics. Today, every major professional sports team has at least an analytics expert on staff, if not a whole department [1]. College teams are increasing their use of analytics as well. Our research goals were to improve the University of Virginia (U. Va.) football team in two ways: recruiting and on-field performance. Our goal of improving the recruiting process led to the development of two tools. First, we created a model that predicts how well an athlete will perform in college based on their high school statistics and demographics. This tool allows coaches to discover lesser ranked athletes who are likely to outperform their rankings. We also further developed an existing model that predicts how likely players are to commit to U. Va. This tool prevents coaches from potentially wasting valuable time and resources on players who are unlikely to commit to U. Va. In order to improve U. Va.'s on-field performance, we created two additional tools. We developed an expected points model based on existing NFL models in an attempt to evaluate the team's performance and identify areas where our play calling was consistently sub-optimal. Finally, we created matchup reports that the coaches can use to scout opposing teams. The expected points model is integrated into these reports to provide a more accurate assessment of the opponent's performance. With this tool, the coaches will be able to spend less time identifying opponents' strengths and weaknesses and more time preparing to exploit them.
Market participants often invoke the concept of discrete state when discussing financial markets. Bull market, bear market, depression, and recession are all terms that map to discrete market states. Mental models of how markets behave in each state and transition between states are then applied to decision-making. Implicit to that approach is the assumption that states are persistent and recurrent over time. This article seeks to formalize notions of discrete market states by proposing a parsimonious and innovative approach to segmenting periods of time into discrete states. The technique is demonstrated and evaluated in a series of case studies.
In this paper, we present a novel method to predict Bitcoin price movement utilizing inverse reinforcement learning (IRL) and agent-based modeling (ABM). Our approach consists of predicting the price through reproducing synthetic yet realistic behaviors of rational agents in a simulated market, instead of estimating relationships between the price and price-related factors. IRL provides a systematic way to find the behavioral rules of each agent from Blockchain data by framing the trading behavior estimation as a problem of recovering motivations from observed behavior and generating rules consistent with these motivations. Once the rules are recovered, an agent-based model creates hypothetical interactions between the recovered behavioral rules, discovering equilibrium prices as emergent features through matching the supply and demand of Bitcoin. One distinct aspect of our approach with ABM is that while conventional approaches manually design individual rules, our agents' rules are channeled from IRL. Our experimental results show that the proposed method can predict short-term market price while outlining overall market trend.
College football programs rely on recruiting to attract high-quality talent, which helps build a team's foundation and ensure success year after year. By conducting systems analysis of the current University of Virginia recruiting process and building upon Walter et al.'s work [1], the Virginia recruiting staff can gain a competitive advantage in the recruiting landscape. Analyzing Virginia's football recruiting and utilizing data analytics could provide the coaching staff with powerful tools to gain such a competitive edge. This study uses a database encompassing over 53,000 football recruits and over 200 predictive attributes to model the four aspects of collegiate football recruiting, as defined by Virginia's football coaches. Specifically, a desirable athlete is defined as one who 1) would succeed on the field at the collegiate level, 2) will meet Virginia's strict academic standards to achieve four years of playing eligibility, 3) fit Virginia football's “gritty” team culture, which is characterized by players who are resilient and able to overcome challenges, and 4) would commit to Virginia if given an offer.
Data analytics has permeated the sports world, but the funds needed to employ data scientists dedicated exclusively to programs at the collegiate level are hard to come by. As a result, Division I football programs are currently not using data analytics and technology to their full potential. This project aims to use technology and data analytics to enhance the performance of the University of Virginia Football program and serves as a continuation of the efforts of previous U.Va. Systems Engineers to bring U.Va.'s program to the technological forefront of college programs. This paper outlines ongoing efforts involving supplementation to analyses currently being done by the program while also introducing new methods to aid the U.Va. coaching staff. Building upon the previously identified “Three Pillars of Data Analytics” for U.Va. Football, this multi-faceted project includes an analysis of how players' practice performance corresponds to in-game performance, the validation of a previously built 4th down decision tool delivered last year as a means to aid in difficult in-game decision making, and the development of an automated tool to classify offensive play formations using computer vision. These efforts aim to help the coaching staff optimize practice schedules, provide them with data-driven decision-making aids to be used in-game, and increase opponent scouting capabilities while also saving the labor required to do so.
Many problems faced by decision makers today involve the management of large scale, complex systems that can be modeled as state-based control problems, specifically discrete Markov decision process (MDP). Typical examples include transportation systems, defense systems, healthcare networks, financial organizations, and general infrastructure problems. In all of these problems, decision makers have difficulty in forecasting the state of their system in the future and capturing the dynamics of the states over time. In this paper, we discuss, via numerous examples, practical experiences in trying to build such models. Much of the literature discusses theoretical issues of solution convergence and algorithm performance; unfortunately, much of this research does not help with the practical business of building an actual MDP model. Thus, numerous books begin with statement of the nature: "given the state space S. . .." A critical question to the practitioner is the creation of this state space "S." We focus on this first step in the MDP modeling process, an often neglected and difficult step, and we discuss the practical implications and issues associated with the state definition, illustrating these issues with numerous examples. This paper is not meant to be a survey of "state-based" applications or MDP applications, but an overview of experiences building many of these models in diverse applications.
This paper expands on previous work that used topic models to characterize transportation research as evidenced by TRB annual meeting papers. This paper advances the previous work by identifying trends over a longer period, exploring the papers in greater depth, and offering a more comprehensive analysis. Almost two decades (1998–2016) of papers were used in the topic model presented here, which identified the themes of the TRB papers and then sorted the documents by these themes. Because the number of accepted papers has increased, most areas of research have also increased in volume over the past 19 years. However, the topic model suggests that transportation research is becoming more holistic and more global. Research on energy and fuel; alternative transportation modes such as carsharing, bicycles, and buses; accessibility and health; and mobile technology is growing much faster than the increase in accepted papers. Additionally, TRB has accepted more papers from international sources, with many papers coming from China. Research on construction and infrastructure, particularly pavements, bridges, pipes, signs and markings, and barriers and guardrails, has increased but is now a smaller proportion of the body of papers accepted by TRB. This change may be attributable to the high cost of the research laboratories and facilities needed to carry out infrastructure research compared with costs for supporting research in some of the rapidly growing research fields that are more information intensive.
Agent-based modeling (ABM) assumes that behavioral rules affecting an agent's states and actions are known. However, discovering these rules is often challenging and requires deep insight about an agent's behaviors. Inverse reinforcement learning (IRL) can complement ABM by providing a systematic way to find behavioral rules from data. IRL frames learning behavioral rules as a problem of recovering motivations from observed behavior and generating rules consistent with these motivations. In this paper, we propose a method to construct an agent-based model directly from data using IRL. We explain each step of the proposed method and describe challenges that may occur during implementation. Our experimental results show that the proposed method can extract rules and construct an agent-based model with rich but concise behavioral rules for agents while still maintaining aggregate-level properties.