.
Interactive physiotherapy and rehabilitation systems require accurate motion assessment and timely feedback to support effective training and user engagement. This study proposes an integrated Human–Computer Interaction (HCI) framework combined with LSTM networks for analyzing sequential 3D joint-trajectory data and providing interactive feedback for motion assessment. The proposed pipeline includes motion data preprocessing and normalization, temporal sequence modeling using LSTM to capture movement dynamics, and an interactive interface for delivering feedback in near real time. We evaluate the framework on motion sequences derived from the Human3.6 M dataset under the described experimental setting. The proposed method achieves up to 91
PurposeWith an explosion of datasets available on the Web, dataset search has gained attention as an emerging research domain. Understanding users' dataset behaviour is imperative for providing effective data discovery services. In this paper, the authors present a study on users' dataset search behaviour through the analysis of search logs from a research data discovery portal.Design/methodology/approachUsing query and session based features, the authors apply cluster analysis to discover distinct user profiles with different search behaviours. One particular behavioural construct of our interest is users' expertise that the authors generate via computing semantic similarity between users' search queries and the title of metadata records in the displayed search results.FindingsThe findings revealed that there are six distinct classes of user behaviours for dataset search, namely; Expert Research, Expert Search, Expert Explore, Novice Research, Novice Search and Novice Explore.Research limitations/implicationsThe user profiles are derived based on analysis of the search log of the research data catalogue in this study. Further research is needed to generalise the user profiles to other dataset search settings. Future research can take on a confirmatory approach to verify these user groups and establish a deeper understanding of their information needs.Practical implicationsThe findings in this paper have implications for designing search systems that tailor search results matching the diverse information needs of different user groups.Originality/valueWe propose for the first time a taxonomy of users for dataset search based on their domain expertise and search behaviour.
In support of open and reproducible research, there has been a rapidly increasing number of datasets made available for research. As the availability of datasets increases, it becomes more important to have quality metadata for discovering and reusing them. Yet, it is a common issue that datasets often lack quality metadata due to limited resources for data curation. Meanwhile, technologies such as artificial intelligence and large language models (LLMs) are progressing rapidly. Recently, systems based on these technologies, such as ChatGPT, have demonstrated promising capabilities for certain data curation tasks. This paper proposes to leverage LLMs for cost-effective annotation of subject metadata through the LLM-based in-context learning. Our method employs GPT-3.5 with prompts designed for annotating subject metadata, demonstrating promising performance in automatic metadata annotation. However, models based on in-context learning cannot acquire discipline-specific rules, resulting in lower performance in several categories. This limitation arises from the limited contextual information available for subject inference. To the best of our knowledge, we are introducing, for the first time, an in-context learning method that harnesses large language models for automated subject metadata annotation.
Persistent identifiers are applied to an ever-increasing variety of research objects, including software, samples, models, people, instruments, grants, and projects, and there is a growing need to apply identifiers at a finer and finer granularity. Unfortunately, the systems developed over two decades ago to manage identifiers and the metadata describing the identified objects no longer scale. Communities working with physical samples have grappled with these three challenges of the increasing volume, variety, and variability of identified objects for many years. To address this dual challenge, the IGSN 2040 project explored how metadata and catalogues for physical samples could be shared at the scale of billions of samples across an ever-growing variety of users and disciplines. In this paper, we focus on how we scale identifiers and their describing metadata to billions of objects and who the actors involved with this system are. Our analysis of these requirements resulted in the definition of a minimum viable product and the design of an architecture that not only addresses the challenges of increasing volume and variety but, more importantly, is easy to implement because it reuses commonly used Web components. Our solution is based on a Web architectural model that utilises Schema.org, JSON-LD, and sitemaps. Applying these commonly used architectural patterns on the internet allows us to not only handle increasing variety but also enable better compliance with the FAIR Guiding Principles.
There is little doubt that we have entered an era where data underpins modern science and research in general. In support of this, numerous infrastructures have been designed and built, ranging from proprietary on-prem systems through to distributed commercial clouds. Such implementations provide a range of functions during the research lifecycle from provisioning and cataloguing data assets through to storing and presenting data to computing platforms. In this paper we analyse the underlying principles of such systems and develop a high-level Research Data Reference Architecture (RDRA). Specifically, we identify eight key features of a RDRA that can guide the design, construction, and procurement of implementations without mandating any domain, approach, technical solution, or product choice. As a result, it allows implementers to make local and commercial decisions while still meeting the core requirements of a research data management platform. The intended audience is teams charged with implementing infrastructure in research organizations.