Answer-set programming (ASP) is a declarative approach to solving combinatorial problems. It lends itself easily to parallel execution. Several ASP solvers are available; we show sample programs in several dialects. The samples include Sudoku puzzles, logic puzzles, Costas arrays, tiling, and Room arrangements. This paper is based on an interactive demonstration presented at the 36th International Workshop on Languages and Compilers for Parallel Computing, Lexington, KY, October 2023.
Code-switching, code-mixing, and, more generally, multilingualism pose technological challenges for language documentation, the sub-discipline of linguistics that deals with the annotation and basic analysis of field recordings and other primary data. We focus here on a case study involving code-mixing in the endangered Koda language, which poses special problems for morphosyntactic analysis. We offer a robust approach to multilingual annotations that involves a combination of the popular open source software FieldWorks Language Explorer (FLEx) with Kratylos, a web-based corpus tool for display and query. Kratylos exposes linguistic data from various formats to powerful regular-expression queries that can exploit tier structure and other aspects of interlinear glossed text. We show how Kratylos can target mixed structures in our FLEx database of Koda that cannot be easily identified within the original FLEx software itself.
The Costas-array problem is a combinatorial constraint-satisfaction problem (CSP) that remains unsolved for many array sizes greater than 30. In order to reduce the time required to solve large instances, we present an Ant Colony Optimization algorithm called m-Dimensional Relative Ant Colony Optimization (mDRACO) for combinatorial CSPs, focusing specifically on the Costas-array problem. This paper introduces the optimizations included in mDRACO, such as map-based association of pheromone with arbitrary-length component sequences and relative path storage. We assess the quality of the resulting mDRACO framework on the Costas-array problem by computing the efficiency of its processor utilization and comparing its run time to that of an ACO framework without the new optimizations. mDRACO gives promising results; it has efficiency greater than 0.5 and reduces time-to-first-solution for the m = 16 Costas-array problem by a factor of over 300.
As data science and machine learning methods are taking on an increasingly important role in the materials research community, there is a need for the development of machine learning software tools that are easy to use (even for nonexperts with no programming ability), provide flexible access to the most important algorithms, and codify best practices of machine learning model development and evaluation. Here, we introduce the Materials Simulation Toolkit for Machine Learning (MAST-ML), an open source Python-based software package designed to broaden and accelerate the use of machine learning in materials science research. MAST-ML provides predefined routines for many input setup, model fitting, and post-analysis tasks, as well as a simple structure for executing a multi-step machine learning model workflow. In this paper, we describe how MAST-ML is used to streamline and accelerate the execution of machine learning problems. We walk through how to acquire and run MAST-ML, demonstrate how to execute different components of a supervised machine learning workflow via a customized input file, and showcase a number of features and analyses conducted automatically during a MAST-ML run. Further, we demonstrate the utility of MAST-ML by showcasing examples of recent materials informatics studies which used MAST-ML to formulate and evaluate various machine learning models for an array of materials applications. Finally, we lay out a vision of how MAST-ML, together with complementary software packages and emerging cyberinfrastructure, can advance the rapidly growing field of materials informatics, with a focus on producing machine learning models easily, reproducibly, and in a manner that facilitates model evolution and improvement in the future.
We describe the use of Kratylos, a web-based corpus query tool, to analyze the morphosyntax of multilingual texts. Among other formats, Kratylos accepts interlinear glossed texts in XML as produced through Fieldworks Explorer (FLEx), the corpus/lexicon building software most commonly employed in language documentation and description. Kratylos facilitates multitier regular expression queries over such XML corpora, which can be used to investigate interactions across languages in a multilingual FLEx corpus, a use case which FLEx itself is not well suited for. We exemplify with a Koda language (cdz) corpus in development that includes a range of borrowing, code mixing and code switching with Bangla.
The last decade has seen great advances in the development of electronic tools for automated interlinearization, corpus creation and lexicon building (e.g. Fieldworks Explorer [FLEx]), as well as tools for creating time-aligned annotations (e.g. ELAN). However, methods for sharing these new data formats online lag far behind. While good options exist for lexical data (e.g. Webonary, Lexique Pro), there is no tool for turning a project created in the FLEx software into an online interlinearized corpus. We present here a tool in development which does precisely that. FLEx databases can be searched using regular expressions and individual lines from a text can be linked to audio and video media. The tool can furthermore bring together linguistic data in diverse formats (from ELAN, Praat, Fieldworks, Toolbox, Shoebox) for a single query and allow for queries over multiple language projects. We discuss the benefits of this program in relation to several ongoing fieldwork projects that are being used to evaluate it. These projects present several interesting challenges. In one, we attempt to create a unified database from several centuries of documentation during which the language showed considerable change. Similarly, in the second project we create a unified database for two lexically, syntactically and phonologically distinct dialects of the same language and show how an interlinearized database facilitates searching across dialects. Finally, in the third project, we show how video data can be integrated into an online FLEx database, a feature which is still lacking in the FLEx software itself. By way of conclusion, we show the audience how to upload their own data (either privately or publicly) and experiment with the tool’s features. Ultimately, the open source program will be available for anyone interested in hosting their own installations.
AbstractThe Papuan language Mian allows us to refine the typology of nominal classification. Mian has two candidate classification systems, differing completely in their formal realization but overlapping considerably in their semantics. To determine whether to analyse Mian as a single system or concurrent systems we adopt a canonical approach. Our criteria – orthogonality of the systems (we give a precise measure), semantic compositionality, morphosyntactic alignment, distribution across parts of speech, exponence, and interaction with other features – point mainly to an analysis as concurrent systems. We thus improve our analysis of Mian and make progress with the typology of nominal classification.
Abstract The Word and Paradigm approach to morphology associates lexemes with tables of surface forms for different morphosyntactic property sets. Researchers express their realizational theories, which show how to derive these surface forms, using formalisms such as Network Morphology and Paradigm Function Morphology. The tables of surface forms also lend themselves to a study of the implicative theories, which infer the realizations in some cells of the inflectional system from the realizations of other cells. There is an art to building realizational theories. First, the theories should be correct, that is, they should generate the right surface forms. Second, they should be elegant, which is much harder to capture, but includes the desiderata of simplicity and expressiveness. Without software to test a realizational theory, it is easy to sacrifice correctness for elegance. Therefore, software that takes a realizational theory and generates surface forms is an essential part of any theorist’s toolbox. Discovering implicative rules that connect the cells in an inflectional system is often quite difficult. Some rules are immediately apparent, but others can be subtle. Software that automatically analyzes an entire table of surface forms for many lexemes can help automate the discovery process. Researchers can use Web-based computerized tools to test their realizational theories and to discover implicative rules.
Abstract We regard the complexity of any inflection-class system as the extent to which the similarities among its inflection classes tend to inhibit motivated inferences about the word forms realizing a paradigm’s cells. We propose ten objective measures of this sort of complexity. We apply these measures in comparing the declensional systems of Latin and Sanskrit, which we represent in a standard format that we call a “plat”; we execute these measurements with an online tool that is freely available for readers to use. We show that the ten measures are not equivalent; together, they show that the declensional systems of Latin and Sanskrit are roughly comparable in complexity. We discuss a number of methodological issues raised by this new approach to typological comparison.
We demonstrate the Binary Identified Contact Domain (BICD) method for predicting binding sites for proteins that starts with amino-acid sequences of several species, finds regions that are evolutionarily conserved, and weights regions by their local hydrophilicity. The measure of conservation takes transversions as significant but transitions as insignificant, in keeping with a theory of original singlet codes. The BICD method correctly predicts several known binding sites for the BRCA2 protein with PABL2 and RAD51 as well as the binding sites of the MDM2-P53 complex.
1. Principal parts 2. Plats 3. A typology of principal-part systems 4. Inflection-class transparency 5. Grammatically enhanced plats 6. Impostors and heteroclites 7. Stems as principal parts 8. The marginal detraction hypothesis 9. Inflection classes, implicative relations and morphological theory 10. Entropy, predictability and predictiveness 11. The complexity of inflection-class systems 12. Sensitivity to plat presentation 13. The Principal-Parts Analyzer.
In this radically new approach to morphological typology, the authors set out new and explicit methods for the typological classification of languages. Drawing on evidence from a diverse range of languages including Chinantec, Dakota, French, Fur, Icelandic, Ngiti and Sanskrit, the authors propose innovative ways of measuring inflectional complexity. Designed to engage graduate students and academic researchers, the book presents opportunities for further investigation. The authors' data sets and the computational tool that they constructed for their analysis are available online, allowing readers to employ them in their own research. Readers can access the online computational tool through www.cambridge.org/stump_finkel.
Rajendra Yavatkar合作论文数System-on-Chip (SoC) Architecture for the Intel Architecture Group at Intel Corporation3
Arkady Zaslavsky合作论文数Caulfield School of IT2
Jon Louis Bentley合作论文数Bell Laboratories2