Egyptian dates are widely used for fixing the chronologies of surrounding countries in the Ancient Near East. But the astronomical basis of Egyptian chronology is shakier than generally assumed. The moon dates of the Middle and New Kingdom are here re-examined with the help of experiences gained from Babylonian astronomical observations. The astronomical basis of the chronology of the New Kingdom is at best ambiguous. The conventional date of Thutmose III's year 1 in 1479 BC agrees with the raw moon dates, but it has been argued by several Egyptologists that those dates should be amended by one day, and then the unique match is 1504 BC. The widely accepted identification of a moon date in year 52 of Ramesses II, which leads to an accession of Ramesses II in 1279 BC, is by no means certain. In my opinion that accession year remains nothing more than one of several possibilities. If one opts for a shortened Horemhab reign, dating Ramesses II to 1290 BC gives a better compromise chronology. But the most convincing astronomical chronology is a long one: Ramesses II in 1315 BC, Thutmose III in 1504 BC. It is favored by Amarna-Hittite synchronisms and a solar eclipse in the time of Mursili II. The main counter-argument is that this chronology is at least 10-15 years higher than what one calculates from the Assyrian King List and the Kassite synchronisms. For the Middle Kingdom on the other hand, among the disputed dates of Sesostris III and Amenemhet III one combination turns out to be reasonably secure: Sesostris III's year 1 in 1873/72 BC and Amenemhet III's 30 years later.
The stochastic behavior of the length of day (LOD) process is analyzed and is modeled within statistical accuracy on a time-scale ranging from weeks to millennia by a three-component model comprising a global Brownian motion process, decadal fluctuations, and a 50-day Madden–Julian oscillation. While the model is intended to be phenomenological, some possible physical models underlying the three components are speculated upon. The model is applied to estimate long-range extrapolation errors. For example, it predicts a standard error of 1 h in the clock-time correction ΔT for extrapolation by 1,500 years from 500 to 2000 BC.
In robustness, as in every area he touched, John Tukey produced hundreds of original ideas, some brilliant, fundamental and lasting, some ephemeral. He presented them in a rambling fashion, in survey articles, in technical reports, in mimeographed draft memoranda and some only in lectures. Along the route, like The Three Princes of Serendip (a favorite fable of his), he would make unexpected discoveries by accident and sagacity. When I started working in robustness in 1961, I much profited from the superb reprint and preprint collection then housed in the coffee room of the Berkeley Statistics Department, containing a fair number of his unpublished papers. Tukey loved words, especially those he had created himself, and his baffling terminology made those papers hard to understand, especially to an outsider like me. The precise meaning of his words sometimes changed over time, and important ideas occasionally got lost in the passage from a preliminary to the final version of a paper.In Tukey's oeuvre, his contributions to robustness are among the least organized, and, regrettably, there is still no corresponding Collected Works volume. I shall try to identify some of Tukey's more important or fertile contributions and to separate them into four categories: conceptual; tools; techniques; and procedures. The separation between the latter three categories admittedly is somewhat fuzzy, but it clarifies matters if we do not throw things such as diagnostic tools, Monte Carlo techniques and estimation procedures all into the same pot. At the same time I shall try to transmit some impression of John's idiosyncrasies and style of interaction. The choice and assessment of the importance of his contributions are mine and sometimes would differ from his own.
. Up to now, interactive statistics seems to have bloomed on a creative mess of contradictions and conflicts. But I believe that if we want to advance beyond the present stage, we must try to overcome those. I shall attempt to identify a few of the more crucial ones, ranging from the basics of interactivity and data graphics to issues of large data and data base management, and to distill specific challenges from them.
Recherche sur les mesures et les decouvertes astronomiques de Babylone ; interpretation et analyse dans l'astronomie lunaire de Babylone du Lunar sixes un ensemble de six intervalles temporels specifiques entre le coucher et le lever du soleil et de la lune, observes aux alentours de la nouvelle lune et de la pleine lune, qui occupe une place essentielle dans les calculs de Babylone.
The three distinct data handling cultures (statistics, data base management and artificial intelligence) fInally show signs of convergence. Whether you name their common area “data analysis” or “knowledge discovery”, the necessary ingredients for success with ever larger data sets are identical: good data, subject areaexpertise, access to technicalknow-how in all three cultures, and a good portion of common sense. Curiously, all three cultures have been trying to avoid common sense and hide its lack behind a smoke-screen of technical formalism. Huge data sets usually are notjust more of the same, they have to be huge because they are heterogeneous, with more internal structure, such that smaller sets would not do. As a consequence, subsamples and techniques based on them, like the bootstrap, may no longer make sense. The complexity of the data regularly forces the data analyst to fashion simple, but problemand data-specific tools from basic building blocks, taken from data base management and numerical mathematics. Scaling-up of algorithms is problematic, computational complexity of many procedures explodes with increasing data size; for nrrsmi-.lD nnn.r.3nt;Annl n,..n+m.;mrr nl",.,‘+hlrmn l.sn.-..%.a "*cLLLAyI.., ~"II"tiIILl"‘Ial n"~wLAJqj a,~"LlLLr‘,m "GC"LI,T unfeasible. The human ability to inspect a dataset, or even only a meaningful of part it, breaks down far below terabyte sixes. I believe that attempts to circumvent this by “automating” some aspects of exploratory analysis are futile. The available success stories suggest that the real function of data mining and KDD is not machine discovery of interesting structures by itself, but targeted extraction and reduction of data to a size and format suitable for human inspection. By necessity, such preprocessing is ad hoc, data specific and driven by working hypotheses based on subject matter expertise and on trial and error. Statistical common sense which traps to avoid, handling of random and systematic errors, and where to stop is more important than specific techniques. The machine assistance we need to step from large to huge sets thus is an integrated computing environment that allows easy improvisation and retooling even with massive data. 1. Copyright Q 1997. American Association for Artificial Intelligence (www.aaai.org). All rights reserved. Introduction Knowledge Discovery in Databases (KDD) and Data Analysis (DA) share a common goal, namely to extract meaning from dam. The only discernible difference is that the former commonly is regarded as machine centered, the latter as centered on statistical techniques and probability. But there are signs of convergence towards a common, human-centered ..:,... -.. l.,.rl. ,:,a,, VlGW WI uulu islUGs. Note the comment by Brachman and Anand (1996, p.38): “Overall, then, we see a clear need for more emphasis on a human-centered, process-oriented analysis of KDD”. One is curiously reminded of Tukey’s (1962) plea, emphasizing the role of human judgment over that of mathematical proof in DA. It seems that in different periods each professional group has been trying to squeeze out human common sense and to hide its lack behind a smoke screen of its own technical formalism. The statistics community appears to be further n,,,.T,v ,x.T l ha .-,“., n mn:r\l;hl *,x1.1 hnn n,,...:,n,,A~,. A, :rl,n cu”Ut; “11 LllFi way, a maJ”l*ly ll”W rum &yulGBLr;u L” LllG *ur;a that DA ought to be a human-centered process, and I hope the Al community will follow suit towards a happy reunion of resources.
Some issues crucial for the analysis of massive datasets are identified: computational complexity, data management questions, heterogeneity of data, customized systems, and some suggestions are offered on how to confront the challenges inherent in those issues.
The development of selected robustness concepts since their inception in the 1960's is sketched and their current status is reviewed.
The three distinct data handling cultures (statistics, data base management and artificial intelligence) fInally show signs of convergence. Whether you name their common area "data analysis" or "knowledge discovery", the necessary ingredients for success with ever larger data sets are identical: good data, subject areaexpertise, access to technicalknow-how in all three cultures, and a good portion of common sense. Curiously, all three cultures have been trying to avoid common sense and hide its lack behind a smoke-screen of technical formalism. Huge data sets usually are notjust more of the same, they have to be huge because they are heterogeneous, with more internal structure, such that smaller sets would not do. As a conse- quence, subsamples and techniques based on them, like the bootstrap, may no longer make sense. The complexity of the data regularly forces the data analyst to fashion simple, but problem- and data-specific tools from basic building blocks, taken from data base management and numerical mathematics. Scaling-up of algorithms is problematic, computational com- plexity of many procedures explodes with increasing data size;
The word “strategy” — literally: the leading of the army“ — inevitably evokes military associations. One is reminded of Clausewitz’ famous treatise ”Vom Kriege“, the first systematic and comprehensive treatment of the subject, published in 1832 after the author’s death. In data analysis, strategy is a relatively recent innovation. A decade ago, in a talk on ”Environments for supporting statistical strategy“ I had quipped that it was difficult to support something which did not exist (Huber 1986). Today, the joke might no longer be appropriate, but we still are far away from a Clausewitz for data analysis.
The three distinct data handling cultures (statistics, data base management and artificial intelligence) finally show signs of convergence. Whether you name their common area "data analysis" or "knowledge discovery", the necessary ingredients for success with ever larger data sets are identical: good data, subject area expertise, access to technical know-how in all three cultures, and a good portion of common sense. Curiously, all three cultures have been trying to avoid common sense and hide its lack behind a smoke-screen of technical formalism. Huge data sets usually are not just more of the same, they have to be huge because they are heterogeneous, with more internal structure, such that smaller sets would not do. As a consequence, subsamples and techniques based on them, like the bootstrap, may no longer make sense. The complexity of the data regularly forces the data analyst to fashion simple, but problem- and data-specific tools from basic building blocks, taken from database management and numerical mathematics. Scaling-up of algorithms is problematic, computational complexity of many procedures explodes with increasing data size; for example, conventional clustering algorithms become unfeasible. The human ability to inspect a data set, or even only a meaningful of part it, breaks down far below terabyte sizes. I believe that attempts to circumvent this by "automating" some aspects of exploratory analysis are futile. The available success stories suggest that the real function of data mining and KDD is not machine discovery of interesting structures by itself, but targeted extraction and reduction of data to a size and format suitable for human inspection. By necessity, such preprocessing is ad hoc , data specific and driven by working hypotheses based on subject matter expertise and on trial and error. Statistical common sense - which traps to avoid, handling of random and systematic errors, and where to stop - is more important than specific techniques. The machine assistance we need to step from large to huge sets thus is an integrated computing environment that allows easy improvisation and retooling even with massive data.
Languages for data analysis and statistics must be able to cover the entire spectrum from improvisation and fast prototyping to the implementation of streamlined, specialized systems for routine analyses. Such languages must not only be interactive but also programmable, and the distinctions between language, operating system, and user interface get blurred. The issues are discussed in the context of natural and computer languages, and of the different types of user interfaces (menu, command language, batch). It is argued that while such languages must have a completely general computing language kernel, they will contain surprisingly few items specific to data analysis—the latter items more properly belong to the “literature” (i.e., the programs) written in the language.
We identify and discuss some of the problems encountered in the data analysis of large data sets, and we suggest some strategies for overcoming them.
In robust estimation, L1-methods serve two main purposes: They provide estimates with minimal bias if the observations are asymmetrically contaminated, and they furnish convenient starting values for estimates based on iterative procedures. In general, such L1-estimates are weighted ones, but there are problems with the choice of weights. It is pointed that the two most popular choices of weights are non-robust: with constant weights, the estimates are sensitive to outliers at influential points, and with Hampel—Krasker—Welsch-type weights they are sensitive to local shifts. Other M-estimates of regression suffer from the same problem. Some recommendations for qualitative improvement are given.