There are several projects in the research community to make the citation data extracted from research papers more re-usable. This paper presents results from the CyrCitEc project to create a publicly available source of open citation content data extracted from PDF papers available at a research information system. To reach this aim the project team has created four outputs: (1) an open source software to parse papers' metadata and full text PDFs; (2) an open service to process papers' PDFs to extract citation data; (3) a dataset of citation data, including citation contexts (currently mostly for papers in Cyrillic); and (4) a visualization tool that provides users insight into the citation data extraction process and gives some control over the citation data parsing quality.
The paper presents first results of the CitEcCyr project funded by RANEPA. The project aims to create a source of open citation data for research papers written in Russian. Compared to existing sources of citation data, CitEcCyr is working to provide the following added values: a) a transparent and distributed architecture of a technology that generates the citation data; b) an openness of all built/used software and created citation data; c) an extended set of citation data sufficient for the citation content analysis; d) services for public control over a quality of the citation data and a citing activity of researchers.
In this paper, we present an approach for analyzing the behavior of editors in the large current awareness service "NEP: New Economics Papers". We processed data from more than 38,000 issues derived from 90 different NEP reports over the past ten years. The aim of our analysis was to gain an inside to the editor behaviour when creating an issue and to look for factors that influence the success of a report. In our study we looked at the following features: average editing time, the average number of papers in an issue and the editor effort measured on presorted issues as relative search length (RSL). We found an average issue size of 12.4 documents per issue. The average editing time is rather low with 14.5 minute. We get to the point that the success of a report is mainly driven by its topic and the number of subscribers, as well as proactive action by the editor to promote the report in her community.
Identifying authorship correctly and efficiently is a difficult problem when the literature is abundant, but poorly recorded. Homonyms are tedious to differentiate. This paper describes how the field of economics has organized itself with respect to author identification. We describe the RePEc project with a special emphasis on the RePEc Author Service. We then discuss how the concept is currently being expanded to the entire scientific body with the AuthorClaim project.
Purpose Applications of information technology have been directly responsible for the increase in productivity of business, government and academic activities. Business and management historians have yet to contribute to better understanding such processes. This paper aims to address this shortcoming through the internal and organisational history of a system for speedy, online distribution of recent additions to the broad literatures on economics and related areas called NEP: New Economic Papers. Design This is a first person account (partly autobiographical) which also includes interviews and the use of archived e-mail correspondence. Findings The advent of the Internet promised a revolutionary change by democratising the social institutions related to the creation and dissemination of academic knowledge. Instead, this story tells how participants slowly but steadily tended to replicate established institutions. Research limitations Researching the impact of the Internet on organizations is a promising topic for historians, for which this might be one case study.Practical implications The development of NEP provides an illustrative example for the kind of new business models that have emerged as the Internet has been used by creative minds to provide existing services in a new way.Social implications This paper provides a story of the NEP project and shows how one person’s drive could generate a broader community of volunteers (constituted by a large number of academics and practitioners who provide critical support for its functioning). We provide details of the social and technological challenges for the construction of the technological platform as well as the evolution of its governance.Originality There is no historiography in business and management history on how to deal with changes in archived material resulting from the application of information and telecommunication technologies. Given the rate of change for events in the third industrial revolution, this article shows is its possible and indeed relevant to document events in the recent past.
We report on the concept, progress and status of the new project Distributed Open Access Reference Citations Services, DOARC . The emphasis is to exploit especially the Open Access documents on Institutional Repositories: analyzing their reference lists for citations, analyzing the full text for research field specific two-word shingles, and use this for powerful author and user tools, one of which can be tested in a demonstrator: plotting the content affinity of scientific papers in dynamically generated graphs. Author-identification is realized by AuthorClaim .
We report on the concept, progress and status of the new project Distributed Open Access Reference Citations Services, DOARC . The emphasis is to exploit especially the Open Access documents on Institutional Repositories: analyzing their reference lists for citations, analyzing the full text for research field specific two-word shingles, and use this for powerful author and user tools, one of which can be tested in a demonstrator: plotting the content affinity of scientific papers in dynamically generated graphs. Author-identification is realized by AuthorClaim .
This article finds an explanation for the subscription to those various fields in the Research Paper in Economics (RePE) e-library, which form the RePEc reports. Subscription to a particular report is regressed on the value as to which papers in that field create as measured by the impact factor of the relevant field journal and the ranking. Subscription has been equated to demand for the fields by analysing demand in cardinal measurement and ordinal measurement. We find that both measurements give rise to demand for digital research data, which is explainable by subsequent value created in the form of publication and citations in high-quality journals for the respective fields.
RePEc (Research Papers in Economics) is the largest academic open access digital library world wide. It references more than 900.000 documents in electronic format. Download of articles and papers it is around 700.000 a month in the last year. In its thirteen years of life, it has become in an indispensable information source of the latest research results in Economics. The RePEc architecture is distributed. More than 1200 institutions collaborate to build a public dataset. Several value added services use the data: alerts about new contents, citation indexing, etc. All contents are re-usable so that any institution may use them for its own community or to build new services for end users. The RePEc business model is based on the same principles as the open source movement: volunteer collaboration through common procedures to add metadata to a shared, open access and re-usable dataset. At a time when the business models of other similar initiatives like arXiv are under review due to sustainability problems, it is interesting analyze the RePEc model. This model has proven strong and sustainable since it is based in a community of users with shared interests. In this paper we introduce RePEc, analyze its business model study its strengths and weaknesses.
In this paper, we discuss the provision of bibliographic data as an extension of the open source concept. Our particular concern is the sustainability of such endeavors. We describe the RePEc (Research Papers in Economics) project, probably the largest ‘open source’ bibliographic database. It demonstrates that open-source bibliographic data collection is sustainable.
Most contributions in this issue are concerned with open source software (OSS) in libraries. Their basic angle is to look at what is being done with OSS in libraries - or what can be done. This contribution takes a broader look. It outlines a number of direct correlations between the functions of libraries and the characteristics of OSS, and by extension, how the principles of OSS can be applied to the distribution of “open libraries” as a future direction for librarianship. Software is nothing but information. The OSS communities create and maintain a bundle of highly structured information for free. What are the implications for the library community? Can they learn something for the open source communities? In other words, I want to look at what can be learned from the OSS software to understand the changing nature of libraries. Libraries are changing dramatically at this time because we are moving from print storage to digital storage and from slow physical transport to fast transport via computer networks. I will start by drawing a parallel between software and libraries. It may appear to be far fetched, but I hope it is nevertheless interesting. To draw the comparison, we have to put software and libraries on an equal footing. A library is most commonly thought of as a service. In our context this description is especially true when we are referring to digital libraries. Normally, software will be thought of as something enabling a service like a digital library. But I would like to look at the software itself as a service. Thus, I will treat both libraries and software as services. Let us start with the software service. Conceptually, a software service can be thought of as three things. First there is something that a user can use. I can open a document in, say, OpenOffice, and I see a bunch of interface elements such as a prompt for keyboard input, some buttons and some icon that moves with the mouse. These interface elements allow me to manipulate a document. In principle, I can imagine another interface to the software service. And often enough, the same piece of software supports slightly different interfaces depending on what computer system it runs on. Second, there is the code that makes the software work. This code is usually some textual data called the “source code.” OSS is software for which the customer who acquires the software also gets the source code. Finally, there are the objects that the software service manipulates. Some may protest that the objects manipulated by the software service are separate from the software itself, but surely, every piece of software is tailored to the objects it manipulates. For example, picture-editing software needs to know the structure of the picture that is its underlying object. If that structure would change, the software would be next to useless, and the software service would be broken. Of course, OSS is not about making manipulated objects freely available. But generally, if many manipulated objects are freely available that's good for the software itself, because it is cheaper to get hold of existing objects to manipulate. Let us turn to the library service. In a similar way as I have explained for the software service there are three elements to a library. There is an interface through which the collection can be accessed. Whether the library is a building or whether the library is a digital collection accessible via a website does matter to the interface. Both types have very different interfaces. What is important here is that the interface can be thought of as a separate component of the library. For example, we can move a physical collection from one building to another. Only the interface changes. Second there is the description of the collection. Like the middle component of the software, this description is the central part of the library. It contains descriptions of the objects held, as well as links between the objects and the users. Finally, there are the objects that the library holds. In a digital library these objects are usably referred to as “full-text files.” In a physical library they are physical books and periodicals. Again, as in the case of the software service, the objects that are manipulated by the library do not necessarily have to be freely available, but it will help the library if they have liberal licensing conditions. Thus we can think of the source code as the core of the software and the description of objects as being the heart of the library. Then we can draw a parallel between open source software and open libraries as shown in Table 1. I have been working on practical open library development. But as sometimes I am asked what I actually do, I have had to do some thinking about the topic. I suspect that I was the first to think of an open library as a freely available collection of descriptions of digital objects that people can reuse and/or change, just like developers can reuse and change open source computer code. At the PEAK meeting at the University of Michigan in March 2000, I presented a paper “RePEc, an Open Library for Economics” archived at http://eprints.rclis.org/archive/00014408/. In September 2004 Michael E.D. Koenig and I presented a paper “From open access to open libraries: Claims and visions for Open Academic Libraries” archived at http://eprints.rclis.org/archive/00002202. These papers have some early thinking about the concept. My concept of an open library is probably best described in these papers. It reflects the creation of freely available digital libraries that are independent of end-user services or any specific usage an end user might make of them. Still, the idea is quite concrete because OSS movement has inspired it. The concept of an open library is very closely aligned to what OSS is about. OSS is really not a project that one organization runs. Rather it is very large set of small-scale projects, many of which achieve great things because they are compatible with others. A bit of history helps. Since the 1980s Richard M. Stallman has called for GNU. GNU stands for “GNU is not UNIX.” It is a free replacement of UNIX. UNIX was a popular operating system. The way the UNIX operating system is built helps the work to replace it. UNIX is not a monolithic system. Instead, it has a lot of components that all work together. Thus GNU project participants did not have to write the whole thing from scratch. Instead, they could start by rewriting utilities such as “ls,” a program that lists file names. The GNU version of “ls” was a drop-in replacement for a standard version in (almost) any UNIX system. GNU versions of such utilities were very often faster and always more full-featured than the standard version. There are lots of ways to write out a list of files. Modern versions of GNU “ls” cover all the most useful ones. Still, I guess that in the early days, most people were skeptical about whether the GNU project could be successfully completed. And it is not really quite complete. But the spirit of GNU lives on, and it lives more strongly than most people ever expected. Free operating systems for computers are a reality. They may not be called GNU systems, but they are free in the sense that Stallman envisioned. Nowadays computers have a lot more functionality than they did in the 1980s. Putting many of these components together on a single computer generates a very complicated system. Let me illustrate this point with a simple example. My operating system of choice is Debian GNU/Linux. The system consists of a set of packages. When I looked at it in May 2008, there were 22456 packages available in my (typical) installation. Each package provides a particular functionality. When I want to add a new functionality to my computer, I add a package, say package A. But packages are not independent. More often than not, when I add a new package, I am told that I have to install a bunch of other packages as well because without these, my desired package A will not run. And there are also other packages that are suggested by package A. I am told that when I run package A, I may also install package B and C that are just friends of package A. Actually, on a technical level things are even more complicated. Each package comes with a version number. Package A version 1.0 may require package B version 2.0 or higher. It may be incompatible with version 2.1 of package B. And so on. You get the idea. Before I bore you with more technical details, let's move away from technology and look at people. Let's look at people who package the software. Let us call them packagers. Most packagers are not the authors of the software they package. They know the software well, and they know the operating system well. They work at the interface between the software and the operating system. They take the software as an input and contribute to the operating system. In doing that task, they may change some aspects of the software. They make these changes because the software can be changed. After all, this feature is what open source is all about. It is open not only for reading but also for writing. Thus it can be adapted to the requirements of the operating system. Such requirements are, for example, that the operating system requires packages to work together. Or, for another example, that a user may change aspects of the software but still want to have it update gracefully when the latest and greatest version of the software comes out at the operating system level. So you start to understand how come all this OSS is available for free. A person who packages a piece of software will spend only a few hours a week on this activity. She can do this, essentially, in her spare time. So the key to making a huge piece of complex information available is to split the task into small bits and assign a volunteer to each bit. This strategy is a key to success. Another key to success is reuse. And reuse in software comes in the form of libraries. All pieces of software rely on libraries. Yes, geeks use libraries too. But their libraries are actually computer files. These special files contain compiled pieces of code that have already been written by somebody else. When I write software for my digital library systems, I use a language called Perl. The code that I write in Perl is a simple text file. But I am not writing all my software from scratch using the commands that Perl provides. Instead I use some structures of Perl code called modules that contain Perl code that has already been written by somebody else to enable common tasks. These modules form a library of code. So modules are one way we can facilitate reuse. In a similar way, Perl itself is written in a language called C, as immortalized in the Beatles song “Write in C.” C code itself relies on C libraries to achieve common tasks such as showing a character on the screen. The maintainers of Perl reuse these libraries. In the same way, when you use a website, the web server, most likely Apache, will answer your requests. The web server software is also written in C and uses the very same libraries that Perl uses to do common tasks, for example reading a file. Can information professionals create and maintain open libraries, just like computer professionals create and maintain freely available software? From what we have learned in the previous paragraph, it should appear infeasible. In the same way that Richard M. Stallman has challenged computing professional to create free software, I challenge information professionals to maintain open libraries. I would not do it had I not already created one, the RePEc open library (see http://repec.org preprint The easing of controls on interest rates has led Ä Interest rates, India, Banks The text is part of a series Working Papers Number 04/17 27 pages 2004-02-13 http://www.imf.org/external/pubs/ft/wp/2004/wp0417.pdf application/pdf Ila Patnaik Ajai Shah The collection in which the paper was published is described in a separate record: series International Monetary Fund Working Papers The IMF, as the publisher of the collection, is described in a separate record. This data is collected by EDIRC, a central service that registers all economics department and research institutions. International Monetary Fund (IMF) publicaffairs@imf.org (202) 623-7000 700 19th Street, N.W., Washington DC 20431 http://www.imf.org/ (202) 623-4661 Washington, District of Columbia (United States) This paper has been claimed by an author to be hers. The author produced this data using the RePEc Author Service. ilapatnaik@gmail.com National Institute of Public Finance and Policy Satsang Vihar Marg New Delhi 110067 INDIA http://openlib.org/home/ila Ila Patnaik Patnaik Ila The paper was mentioned in the NEP: New Economics Papers report on financial markets, October 22, 2005. This data is contributed by that service. Carolina Valiente 1130233735 There is also a different version of the last set of data that contains all papers in the report with as full descriptions as were available at the time of the report. This latter set of data is important for internal housekeeping at NEP and to feed the statistical learning procedures that help editors compose report issues. In addition, reports also have collection metadata attached to them, which shows the current editor of the report. The record above, which is specific to an issue of the report, shows data about the editor who prepared the issue in which the paper was published, but this editor is not the current editor. There are powerful obstacles to achieving open libraries. I have three for you here. First, there is technical incompetence; then, there are two more sophisticated problems that I call the “myth of industry” and the “myth of the full text.” Let me elaborate on the three obstacles in turn. Technical incompetence is a huge problem. Unicode, XML and its related technologies such as XML Schema and XSTL, CSS, SQL, OAI-PMH and OAI-ORE, operating system skills, basic knowledge of networking, and above all, knowledge of a scripting language such as Perl or PHP - it all adds up to a large body of knowledge. While it is not required that every digital library builder have a deep knowledge of each of these areas, a deep understanding of at least a few of them, as well as having the programming skills, is required. Without this foundation, we can't get started. Usually none of these technologies is taught in library schools. For years, I have been battling to introduce at least a small part of this body into the curriculum of my school, without much success. As a result the average library school graduate has almost no chance of getting involved in digital library building. One may argue that this work can be left to technical staff and that library staff only need to design the system. To characterize how absurd this idea is, I use the analogy of a person who wants to be a singer but has no voice. You can't turn to this person and say, “OK. No problem. You can't sing, but you can just imagine how a piece should be sung, and somebody else will sing for you.” Without having studied the technical underpinnings, library staff lack the analytical reasoning skills that are required even to get started with the design of new systems. All they will be able to say is, “Oh, it should be user friendly.” As a result, innovation in libraries is stifled. There is a tendency to contract out all developments that involve digital information skills. As I wrote in a mailing list recently, “Libraries are outsourcing to their death.” A second problem is what I call the “myth of industry.” It is a term that I coined myself. It is the tendency of people involved in digital library work to protect their work. They put up obstacles to its reuse. Their idea is that they have built the data, and therefore they want to keep tight control over its usage. As a result you have to ask them to get a copy of the data, and more often than not, the answer is no. For example, it has not been possible for me to get a copy of the Astrophysical Data Service to use in the AuthorClaim registration service. What industry mythers do not understand is that by giving the structured data away freely, they encourage reuse of the data. It means that their own contributors have better incentives to contribute data since it is more widely used. In other words, by giving away their structured data, open access publishers and digital library builders add value to their collections. It will still take a long time until this point sinks in. Finally, there is the worship of the full text. There is too much emphasis in libraries on the problem of users reaching full text - where I take a wide view of what “full text” actually is. Many people place the full text of resources at the center of their collection development. If the full text is really a textual object, and if it is freely available as it should be, it can be quite easily indexed by a full-text engine, say Google. Thus textual metadata attached (in some form) to the full-text is not really important. The same textual string can also be found in the full text. So people put documents on a website, have them indexed by Google and say “That's it, I am done.” This approach works if the text is an announcement of your next birthday party, but it appears insufficient when we deal with important documents such as scientific papers, legal codes, technical documentation or works of art. In these fields we are typically not only interested in getting access to the full text, but, in fact, we are also interested in the links among these object. For example, in patent data, we are interested in such things as citation links between patents, who applied for the patent and whether the patent was approved. In academic work we need to know who the author is. In preservation we need to know what general class a full-text object belongs to so that we can reach a decision about whether and - if yes - how to preserve it. All these concerns require registries. And these registries have to be compiled partly at least by hand. This registry creation is the job of digital librarians. Digital librarians are required to set up registries, to monitor their contents and, sometimes, even to populate them. Registries make a digital library more than just a collection of http-available computer files. But work on registries cannot progress unless more people realize their importance. Thus, in digital libraries, the full text should not be considered of central importance. Rather, it should be considered to be a metadata attribute. Libraries traditionally have been working with non-free information. They have argued that resources should be pooled to purchase access to such information for community members. Their promotion of free information has been hypocritical. They have advocated free access to information as long as it requires paying libraries to provide it. The most important trend libraries are facing is the increase of free access information resources. Nowhere is this more obvious than on the web. More and more serious information is being made available for free on websites. Project Gutenberg was an early starter. Many newspapers, for example, have been building websites and offer much of their content for free on these sites. Many institutions offer important information about themselves on websites. Encyclopedic knowledge is more widely available than ever thanks to Wikipedia. Generally, we see societies moving from an economy of information to an economy of attention. In the economy of information, information is rare and attention is plentiful. In the economy of attention, it is the opposite. The fact that we now have a freely available computer operating system is a part of the attention-economy trend. So far the library sector is stuck in the economy-of-information track. It will wither if it does not get out of there. Libraries have the opportunity to participate in the creation of open libraries that provide structured information on behalf of community members for free reuse by others, which can be a value-added business model for them. Building open libraries requires technical skill current librarians generally don't have. It requires a business sense they have problems perceiving. And it requires a change in purpose that they are slow to accept. Therefore, I am not optimistic about the future of the formal library sector. But, of course, open libraries that are modeled after the open source movement are here to stay. Some of this paper was written while I enjoyed the hospitality of Siberian Federal University www.sfu-kras.ru/. I am grateful to Eric Lease Morgan for comments that have helped to improve the paper, including the suggestion to add the metadata example.
AbstractAn open access community is a digital repository or an online community where scientific information and communication are free to the public through computing technologies (Hanauske, M., et al 2007; Hubbard, C., et al, 2005). Open access community provides a new way for knowledge sharing and knowledge management. It takes advantage of collective expertise by providing a repository for research papers and research data that are scattered or take a long time to be published. The panel will discuss experiences and challenges people face in various open access communities. Particularly, we will discuss the following issues: How did each community or repository achieve the functions of “organize” and “share” among people having a common interest in the community? How long did it take to launch and establish an open‐access community? What impact of such an open‐access community / repository has on people's interaction with information? Impact on fee‐based digital libraries or traditional libraries? What the tradeoffs are between opened vs. controlled? How well do they address privacy issues? How well is current open access community/ repository meeting human needs, and what should future technology research and development involve to better meet user needs?
This paper studies the main characteristics of the citation indexes currently developed in Spain. The paper compares the impact factors offered by Spanish citation indexes with the impact factor of Spanish journals also collected by the JCRs of the ISI (SCI and SSCI) over a five-year period (2001–2005). Spanish journals published in English have higher impact factor scores in the JCR databases of the ISI than in Spanish citation indexes.
This is a personal introduction to the relationship between digital libraries and networks. I recall the way I came to the subject of networks. I describe the way that I have tried to harness networks for digital library building. And I point out some of the difficulties in networking digital objects and their descriptions in digital libraries.
Citations to 200 top downloaded papers at RePEc, a digital library in economics, were obtained from SSCI and Google Scholar respectively to address questions relating to downloads and their corresponding citations. This study finds that top downloaded documents are used in various degrees when citation is regarded as an indicator of usage. The results also show that a single downloaded paper selected for this study on average receives twice as many citations from Google Scholar as that from SSCI although the latter has been established much earlier in time. According to the coefficients computed, downloads appear having a moderate relationship with citations. However, other measures such as the download-citation ratio indicate a stronger connection between the two. While an author's reputation positively affects both download and citation frequencies, other factors (e.g., targeted readers and subject content) seem in play differently for the documents that are repeatedly downloaded or cited. The study suggests that an infrastructure which encourages downloading at digital libraries could lead to higher usage of their resources.
“C’EST L’HOMME QUE JE SUIS QUI ME REND MISANTHROPE”, noted Jules Renard. Yes, I don’t care much for the concepts of Web 2.0 and Library 2.0 and other buzz words a la mode. Call me a dinosaur, but I don’t blog and don’t read blogs. But even cynics have to admit that the Internet continues to transform information services like nothing ever before. Some innovations will stay, others will be gone in a few years.
AbstractRePEc (Research Papers in Economics) has been conceived and developed to promote scholarly communication and to enhance the dissemination of research findings in the field of economics. RePEc offers the RePEc Author Service (RAS) where economics authors can claim authorship of the research papers that are described in RePEc archives. The data from this service forms a high‐quality authorship database. We investigate the structure of research collaborations within RePEc by applying social network analysis to the co‐authorship network formed by the RAS registrants. We perform a component size analysis and calculate centrality metrics. Our findings imply that the RAS registrant population is made up of highly active academics that are well connected to each other. In addition, RAS registrants appear to have a broad range of coauthors, with most individuals having only a few coauthors, whereas a few have many. We compare and contrast results from a number of recent studies of similar scope on co‐authorship networks.
Herbert Van De Sompel合作论文数Digital Library Research and Prototyping Team ; Los Alamos National Laboratory;Research Library4
Kurt Maly合作论文数Department of Computer Science, Old Dominion University3
Tim Brody合作论文数School of Electronics, University of Southampton2