
Following longer standing evidence that family businesses innovate differently than non-family firms, we investigate why and how family businesses may manage Intellectual Property (IP) in specific ways. We employ a qualitative research approach interviewing managers from nine family businesses as well as three external IP experts. Drawing on the perspective of socio-emotional wealth (SEW) theory as an interpretive lens, we find evidence that the long-term orientation of family businesses – which seek to preserve the firm for future generations – entails that IP protection is often seen as an investment in future competitiveness and sustainability for the next generation. For such family businesses, IP management is not just a defensive mechanism, but a strategic tool that supports specifically long-term growth, ensures sustainability, and safeguards the family's heritage across generations. Owning families of these firms and their values imprint IP management practices. The specific characteristics of these family firms may result in more cautious and long-term innovation and IP use, the optimization of products and processes as well as a high priority of quality, reflecting the high relevance of reputational concerns as well as close binding ties with stakeholders. These close relationships also extend to the employees who play a pivotal role in IP management.
The key to success in automating prior art search in patent research using artificial intelligence (AI) lies in developing large datasets for machine learning (ML) and ensuring their availability. This work is dedicated to providing a comprehensive solution to the problem of creating infrastructure for research in this field, including datasets and tools for calculating search quality criteria.The paper discusses the concept of semantic clusters of patent documents that determine the state of the art in a given subject, as proposed by the authors. A definition of such semantic clusters is also provided.Prior art search is presented as the task of identifying elements within a semantic cluster of patent documents in the subject area specified by the document under consideration.A generator of user-configurable datasets for ML, based on collections of U.S. and Russian patent documents, is described. The dataset generator creates a database of links to documents in semantic clusters. Then, based on user-defined parameters, it forms a dataset of semantic clusters in JSON format for ML.A collection of publicly available patent documents was created. The collection contains 14 million semantic clusters of US patent documents and 1 million clusters of Russian patent documents.To evaluate ML outcomes, it is proposed to calculate search quality scores that account for semantic clusters of the documents being searched. To automate the evaluation process, the paper describes a utility developed by the authors for assessing the quality of prior art document search.
The synergy between industrial and innovation chains—termed dual-chain synergy—is critical for high-quality regional economic development, yet operationalizing this integration presents significant governance challenges. To overcome the limitations of traditional, heuristic policy planning, this study introduces a multi-methodological “PPS” framework for public policy analysis, wherein a selectively deployed Large Language Model (LLM) augments the TRIZ-based industrial policy component, while innovation and human resource policies are addressed through complementary structural and network-analytic tools. Using the Qin Chuang Yuan innovation platform in Shaanxi Province as a case study, this research demonstrates the framework's efficacy as a verified decision-support mechanism. Industrial strategies and human resource strategies are generated after PSS reasoning that could be applied for the government of Shaanxi Province. Ultimately, this study transcends conventional qualitative reviews, providing policymakers with a reproducible, computationally augmented toolkit to shift regional governance from subjective administration to an objective, evidence-based digital ecosystem.
Patent documents are expected to disclose technical information and delineate the scope of exclusive rights. However, the extent to which their structure facilitates the subsequent identification and use of patent information remains insufficiently understood. This study examines how two structural features of patent documents—the length of independent claims and the length of technical specifications—are associated with a citation-based indicator related to patent information use. Using 765,290 Chinese patents granted between 2005 and 2012, we use non-self forward citations as an empirical proxy for the subsequent visibility and use of patent documents. We interpret this measure cautiously: forward citations do not directly measure patent information quality, but they provide an observable signal that a patent has been identified and used in later patenting activity. The results show an inverted U-shaped association between both independent claim length and technical specification length, on the one hand, and the citation-based indicator, on the other. The number of dependent claims strengthens the inverted U-shaped association for independent claim length, whereas backward citations flatten the association for technical specification length. These findings suggest that both overly concise and overly lengthy patent documents may be less conducive to subsequent information use. The study contributes to research on patent information by providing large-scale empirical evidence on how document structure relates to patent information usability, while also highlighting the need for greater methodological transparency in patent-text extraction and analysis.
With the rapid growth of intellectual property and standardization, standard-essential patents (SEPs) are gaining global attention. Accurately and swiftly predicting SEPs is becoming a key indicator of a country's international influence. However, current research faces challenges: existing methods underutilize diverse patent features, cannot predict patent standardization time, and lack suitable datasets for dynamic SEP prediction. To address these issues, this paper proposes a dynamic SEP prediction model. Given a patent, we first extract its static features, then divide multiple time slots of varying lengths to capture dynamic features. We employ multiple Transformer models to dynamically predict whether the patent will become a SEP in the coming years. Additionally, we construct new datasets specifically for patent standardization time prediction. Comparative experiments show our model significantly outperforms baselines. Ablation studies further reveal the impact of different patent feature types on dynamic SEP prediction.
Accelerating the commercialization of renewable energy technologies is essential for closing the global climate investment gap. However, traditional models of patent valuation, which remain highly dependent on citation-based proxies, frequently fail to capture real-time market signals. To overcome this limitation, this study proposes a hybrid multimodal framework that integrates transformer-based semantic embeddings (BERT and ELECTRA) with gradient boosting algorithms (XGBoost, LightGBM, and CatBoost) to predict patent assignment frequency-a robust, behaviorally grounded indicator of actual commercialization activity. Based on a dataset of 33,570 renewable energy patents (CPC Y02E) issued by the USPTO between 2014 and 2023, we operationalized a predictive pipeline that fuses unstructured textual data (abstracts and full claims) with structured and relational metadata, including inventor/assignee identities and backward citations. The empirical results demonstrate that the hybrid architecture significantly outperforms single-modality baselines, achieving an R2 score exceeding 0.95 and an RMSE as low as 0.62. A pivotal finding is the identification of a "performance crossover": while ELECTRA combined with LightGBM exhibits superior stability across broad, heterogeneous datasets, BERT combined with XGBoost proves significantly more effective in high-selectivity regimes (>= 8 transfers), where a deep contextual understanding of complex technical claims is paramount. By bridging the methodological gap between deep semantic text extraction and structured market signals, this work provides policymakers, investors, and R&D managers with a scalable, data-driven mechanism for identifying high-potential green technologies.
The biosynthesis of silver nanoparticles (AgNPs) has emerged as a sustainable alternative to conventional physicochemical routes, employing renewable biological sources and reducing toxic agents. This study conducts a large-scale technological and scientific prospection of green AgNP synthesis, based on 373 patent families and 12,529 scientific articles indexed up to 2025. S-curve modeling indicates that scientific production maintains active growth (K = 18,000; r = 0.30), approaching 1600 annual publications in 2025, while patent filings enter technological consolidation (K = 440; r = 0.32). In articles, plant extract-based routes predominate (73%), with diverse phenolic compounds as reducing agents, notably Azadirachta indica (9%) and Ocimum sanctum (5%). In patents, more standardizable sources prevail, such as microorganisms (22%), isolated biomolecules, and hybrid systems (17%), with emphasis on Fusarium oxysporum (9%) and Bacillus subtilis (9%). China represented the leading patent publication jurisdiction (60%; CAGR = 18.2%), while India concentrated scientific production (CAGR = 30.5%). Universities represented 59% of applicants. Synthesis parameters converge to consolidated ranges: pH 6.0-10, Ag+ concentration 1-3 mM, and diameter 10-100 nm. Articles employ longer reaction times (1-24 h, 59%) with advanced techniques, whereas patents prioritize accelerated kinetics (<1 h, 57%) and integrated routes. The results evidence a field in transition, with academic experimentation paving the way for industrial consolidation.
Inventive step (non-obviousness) is widely regarded as the most resource-intensive and unpredictable requirement in patent examination. While patent offices provide published detailed guidelines, inventive-step determinations remain vulnerable to interpretive discretion, resulting in significant inconsistency across examiners and jurisdictions. Such instability increases prosecution costs for applicants and reduces predictability in innovation-based business strategy.This paper proposes a structured reasoning framework, referred to as the pair formulas technique, which operationalizes inventive-step examination through two parameters: (i) the technical difference between a claimed invention and the closest prior art, and (ii) the benefit (effect) arising from that difference. This framework is not intended to fully reproduce actual inventive-step practice in all doctrinal and practical respects. Rather, it isolates a narrower analytical core that seeks to make the conventional “motivated to modify” inquiry more transparent, evidence-centered, and repeatable.To explore feasibility, we conducted prompt-based experiments using an off-the-shelf large language model (LLM)-based chatbot configured with strict procedural rules that emulate examiner reasoning. Because the chatbot was not fine-tuned for patent examination, perfect doctrinal compliance was not assumed at the outset. Instead, the experiments examined whether prompt-based structure, together with human verification and limited dialogue-based correction, could organize inventive-step reasoning into a more transparent and auditable process.The results suggest that prompt-based AI systems may serve as practical tools for examiners and applicants by standardizing reasoning steps, and reducing unsupported variability. At the same time, the experiments show that human verification remains necessary, particularly where element mappings or unsupported interpretive drift may arise.
Species of the genus Ganoderma have been used in traditional medicine for centuries and is now a strategic biotechnological resource. Despite its importance, no comprehensive patent analysis has mapped its technological development. This study aimed to clarify the global status, maturity, and innovation frontiers, providing a basis for future advancements. Patent data up to 2025 were analyzed, resulting in 17,920 patent families, using lifecycle modeling, mapping, profiling, and IPC domain analysis. It also compared keyword extraction methods (TF-IDF, KeyBERT) with a Large Language Model (Meta-Llama 3.2-1B), which outperformed traditional methods in generating domain-specific keywords. Patent activity had three phases: emergence (1979-2009), exponential growth (2000-2018), and maturity (2018-2026), with saturation around 2035. China led patenting, followed by Korea, Japan, and the US. Most patents were held by private companies, with academic institutions contributing R&D. IPC analysis indicates a transition from formulation-based technologies (A61K, A61P, A23L), which peaked between 2010 and 2015, toward sustained growth in biotechnological and processing domains (C12N, C12R, B01D, C08B) from 2014 onward. The findings show a mature yet evolving ecosystem, with Ganoderma mushrooms technologies moving from traditional pharmacology to bioprocess and bioeconomy applications.
Traditional patent classification systems use broad labels that are often too imprecise or ill-adapted to distinguish green from circular innovations, particularly in material-intensive sectors with long-standing waste management issues such as tyres. Existing taxonomies frequently conflate resource loop closure (circularity) with broader environmental mitigation (green), resulting in conceptual and practical ambiguity. To address these limitations, we propose a novel method that integrates a transformer-based NLP model (DistilBERT) within an operational framework based on the 3R strategies — reuse, recycling, and recovery. This approach systematically identifies circular patents, revealing patterns and overlaps beyond traditional classifications. Applying our method to 59,298 tyre-sector patent applications filed at the EPO, USPTO, and WIPO, we find that many circular innovations are not captured by previous classifications. Our results demonstrate that transformer-based NLP can provide a scalable and empirically validated approach to delineate the landscape of circular innovation, clarify the boundary between green and circular patents, provide a detailed breakdown of circular patents by 3R-strategies, and derive implications for firms, policymakers, and researchers seeking to monitor the transition to circularity.
This research note analyzes forward citations for patents granted to the three largest U.S. National Laboratories: Los Alamos (LANL), Sandia (SNL), and Lawrence Livermore (LLNL). Using USPTO data from 2005 to 2021, we examine the rarity of high-impact inventions within the federal research system. Grounded in March's (1991) exploration-exploitation framework, we define an Innovation Lottery model where mission-driven R&D is characterized by extreme uncertainty and stochastic breakthroughs. Our results reveal a consistent power-law distribution across all three labs, with scaling exponents (alpha) ranging from 1.88 to 3.02. Sandia (SNL) exhibits the most heavy-tailed distribution (alpha = 1.88), suggesting a robust alignment with dual-use commercial cycles, while LLNL's rapid decay (alpha = 3.02) reflects the friction of highly specialized Big Science. We find that up to 67.7% of patents receive only a single citation, representing the necessary search costs of discovery. Conversely, a tiny fraction of Black Swan outliers accounts for the vast majority of technological spillovers. Aligning with May 2025 GAO findings regarding patent quality, we argue that volume-based KPIs are fundamentally misleading. We propose reforming technology transfer through tiered CRADA structures and successbased royalty recapture to better optimize the head of the impact distribution.
Manual patent classification is labor-intensive and time-consuming. With the rapid growth of patent data and labels, traditional approaches struggle to handle large-scale classification. Existing automated methods partially address this issue, but most focus only on patent text semantics and ignore hierarchical label information. This limits models from fully exploiting label relationships, reducing both accuracy and generalization. In this study, a contrastive learning-based multilabel classification model, named CLMCM, is proposed, which incorporates the global and local hierarchical information into a patent text encoder to accomplish the classification task. Firstly, a hierarchical label-aware embedding module is deployed to exploit the global semantic information to generate label-wise patent text vectors. Secondly, a hierarchical adaptive label correlation learning module is designed to adaptively learn the semantic correlations around the label codes within the hierarchical taxonomy tree, and then aggregate the horizontal and vertical information to capture the local relationship of the labels. Finally, experimental evaluations on a real-world patent dataset demonstrate the superiority of the proposed method over existing multilabel classification methods, achieving 1.6% improvement in Mi-F1 score.
This study investigates statistical differences between human-written patent abstracts and those generated by ChatGPT under controlled conditions. We analyze 500 patent abstracts from Taiwanese applications filed in 2020, a dataset selected to ensure temporal consistency and to minimize potential influence from AI-assisted writing tools. Given that patent abstracts emphasize essential technical features and follow relatively standardized, claim-oriented language, they provide a suitable setting for controlled stylistic analysis.Using restricted inputs (patent titles or original abstracts), we generate corresponding texts with ChatGPT and examine distributional differences via exploratory and confirmatory statistical analysis. The proposed framework, inspired by ecological diversity, models words and phrases as species and employs diversity-based metrics—such as standardized type-token ratio and entropy—as explanatory variables. Experimental results show that a small set of interpretable statistical features can achieve classification performance comparable to a BERT-based model, while requiring substantially fewer variables and lower computational cost. Importantly, the observed differences reflect distributional characteristics under constrained generation conditions rather than universal distinctions between human and AI-generated text.This study provides an interpretable, domain-specific baseline for understanding statistical differences between human and AI-generated language, highlighting the role of context, domain, and input constraints in shaping stylistic patterns.
Technological opportunity identification (TOI) has become increasingly challenging as innovation cycles shorten and technology recombination accelerates across domains. Existing TOI studies often treat semantic content and structural relations separately, rely on shallow keyword co-occurrence, and rarely provide rigorous retrospective validation. This study proposes a patent-based TOI framework that couples deep semantic representations with network topology to detect early signals of cross-functional convergence. First, patent titles and abstracts are encoded using Sentence-BERT to obtain dense semantic embeddings and are clustered via K-means to form interpretable functional modules, which are labeled using TF-IDF keywords. Second, a k-nearest-neighbor semantic similarity network is constructed and partitioned by Louvain community detection to capture structural communities reflecting actual technology recombination patterns. Third, we introduce semantic-structural coupling entropy, an information-theoretic measure that quantifies the heterogeneity of functional modules within each structural community; high-entropy communities are operationalized as potential opportunity regions. Finally, a time-sliced backtesting design links coupling entropy measured in an identification window to the subsequent emergence and diffusion of novel module pairs in a validation window, providing retrospective evidence for predictive validity. An empirical study on global unmanned surface vehicle (USV) patents demonstrates that high-entropy communities concentrate system-integration trajectories and yield a higher intensity and diffusion of novel cross-module combinations, which can be translated into actionable integration schemes. The proposed approach offers a scalable and reproducible tool to prioritize technology opportunities and supports R&D decision-making under increasing technological complexity.
Patent intelligence is increasingly constrained by the volume, technical density, and linguistic variability of patent documents, which limit the feasibility of producing high-granularity, analyst-facing outputs, such as matrices, trajectories, and comparative maps, using conventional keyword- or bibliometric-based methods. This study addresses this feasibility gap by proposing an Artificial Intelligence (AI)-based framework designed to stabilize patent content for scalable strategic analysis. Methodologically, the framework transforms patent documents into structured semantic summaries optimized for large language model processing and uses these representations to support AI-driven clustering and multidimensional classification. A retrieval-augmented generation (RAG) architecture ensures that summaries and classifications remain grounded in the underlying patent corpus, enhancing reliability and transparency. The framework is demonstrated through a large-scale robotics case study focusing on Siemens, Toshiba, and Mitsubishi. Results show improved recall and precision compared with keyword-based searches for semantically ambiguous categories, illustrated through a leggedrobot benchmark. The approach enables the systematic construction of technology-application matrices that reveal competitive positioning, white spaces, and cross-domain innovation patterns. It further supports cell-level deep dives via grounded clustering and evolutionary timelines, as well as technology-acceleration analysis and normalized portfolio benchmarking. An additional application to European funded-project data highlights misalignments between corporate patenting trajectories and public R&D investment priorities. Overall, the study demonstrates that AI-based semantic stabilization provides a robust and scalable foundation for advanced patent intelligence in complex technological domains.