In this article, we present a novel approach to quantitative evaluation of a model for parsing web pages as visual images, intended to provide improvements for users with assistive needs (cognitive or visual deficits, enabling decluttering or zooming and supporting more effective screen reader output). This segmentation-classification pipeline is tested in stages: We first discuss the validation of the segmentation algorithm, showing that our approach produces automated segmentations that are very similar to those produced by real users when making use of a drawing interface to designate edges and regions. We also examine the properties of these ground truth segmentations produced under different conditions. We then describe our Hidden Markov tree approach for classification and present results which serve provide important validation for this model. The analysis is set against effective choices for dataset and pruning options, measured with respect to manual ground truth labelling of regions. In all, we offer a detailed quantitative validation (focused on complex news pages) of a fully pipelined approach for interpreting web pages as visual images, an approach which enables important advances for users with assistive needs.
In this paper we present a mathematical model of the Empirical Mode Decomposition (EMD). Although EMD is a powerful tool for signal processing, the algorithm itself lacks an appropriate theoretical basis. The interpolation and iteration processes involved in the EMD method have been obstacles for mathematical modelling. Here, we propose a novel forward heat equation approach to represent the mean envelope and sifting process. This new model can provide a better mathematical analysis of classical EMD as well as identifying its limitations. Our approach achieves a better performance for a "mode-mixing" signal as compared to the classical EMD approach and is more robust to noise. Furthermore, we discuss the ability of EMD to separate signals and possible improvements by adjusting parameters.
We propose a new method for learning filters for the 2D discrete wavelet transform. We extend our previous work on the 1D wavelet transform in order to process images. We show that the 2D wavelet transform can be represented as a modified convolutional neural network (CNN). Doing so allows us to learn wavelet filters from data by gradient descent. Our learned wavelets are similar to traditional wavelets which are typically derived using Fourier methods. For filter comparison, we make use of a cosine measure under all filter rotations. The learned wavelets are able to capture the structure of the training data. Furthermore, we can generate images from our model in order to evaluate the filters. The main findings of this work is that wavelet functions can arise naturally from data, without the need for Fourier methods. Our model requires relatively few parameters compared to traditional CNNs, and is easily incorporated into neural network frameworks.
The wavelet transform has seen success when incorporated into neural network architectures, such as in wavelet scattering networks. More recently, it has been shown that the dual-tree complex wavelet transform can provide better representations than the standard transform. With this in mind, we extend our previous method for learning filters for the 1D and 2D wavelet transforms into the dual-tree domain. We show that with few modifications to our original model, we can learn directional filters that leverage the properties of the dual-tree wavelet transform.
We propose a novel diffusion-based, empirical mode decomposition (EMD) algorithm for image analysis. Although EMD has been a powerful tool in signal processing, its algorithmic nature has made it difficult to analyze theoretically. For example, many EMD procedures rely on the location of local maxima and minima of a signal followed by interpolation to find upper and lower envelope curves which are then used to extract a “mean curve” of a signal. These operations are not only sensitive to noise and error but they also present difficulties for a mathematical analysis of EMD. Two-dimensional extensions of the EMD algorithm also suffer from these difficulties. Our PDEs-based approach replaces the above procedures by simply using the diffusion equation to construct the mean curve (surface) of a signal (image). This procedure also simplifies the mathematical analysis. Numerical experiments for synthetic and real images are presented. Simulation results demonstrate that our algorithm can outperform the standard two-dimensional EMD algorithms as well as requiring much less computation time.
In this work we propose a method for learning wavelet filters directly from data. We accomplish this by framing the discrete wavelet transform as a modified convolutional neural network. We introduce an autoencoder wavelet transform network that is trained using gradient descent. We show that the model is capable of learning structured wavelet filters from synthetic and real data. The learned wavelets are shown to be similar to traditional wavelets that are derived using Fourier methods. Our method is simple to implement and easily incorporated into neural network architectures. A major advantage to our model is that we can learn from raw audio data.
In this paper we introduce an edge-based segmentation algorithm designed for web pages. We consider each web page as an image and perform segmentation as the initial stage of a planned parsing system that will also include region classification. The motivation for our work is to enable improved online experiences for users with assistive needs (serving as the back-end process for such front-end tasks as zooming and decluttering the image being presented to those with visual or cognitive challenges, or producing less unwieldy output from screenreaders). Our focus is therefore on the interpretation of a class of man-made images (where web pages consist of one particular set of these images which have important constraints that assist in performing the processing). After clarifying some comparisons with an earlier model of ours, we show validation for our method. Following this, we briefly discuss the contribution for the field of computer vision, offering a contrast with current work in segmentation focused on the processing of natural images.
In this paper we present an overview of our proposed algorithms for classifying regions of web pages based on content and visual properties. We show how hidden Markov trees may be effective for the classification and how this may end up offering improved experiences to users who are trying to view webpages.
•We use a novel vision-based method to analyze the layout of a web page.•Our method produces a hierarchical segmentation of the page reflecting its structure.•Vision-based methods are not sensitive to implementation language or complexity.•The visual presentation of a page provides rich information about semantic structure.•This structure can help create modified presentations for users with assistive needs.
Extended 6 Transistors (6T) SRAM (Static Random-Access Memory) characterization is used to measure degradation while separating intrinsic from extrinsic yield and accounting for yield assessment challenges such as voltage drop and measurement variability. Separation of extrinsic yield pre- and post-stress reveals weak yield fixes and reduces HTOL (High Temperature Operating Life) failure risk.
We present a spatial-domain method for reconstructing a three-dimensional density distribution from one or more projections (images formed by integration of density along lines of sight) and using the three-dimensional reconstruction to explain features of the two-dimensional images. The advantages of our proposed method are that it degrades gracefully down to a single image, that it uses linear equations and constraints (allowing the use of convex optimization), that it is amenable to three-dimensional structural biases, and that ambiguity can be expressed precisely (it is possible to "know what we don't know"). Previously described methods have some, but not all, of these properties.
With the increasingly rich display of media on the Internet, screen reading technology that mainly considers website source code can become ineffective. We aim to present a solution that remains robust in the face of dynamically displayed web content, regardless of the underlying web framework. To do this, we consider techniques used in computer vision to determine semantic information about the web pages. We consider existing screen reading technologies to see where such techniques can help, and discuss our analytical model to show how this approach can benefit low vision users.
We present a spatial-domain method for the reconstruction of a three-dimensional density distribution from one or more projections (images formed by integration of density along lines of sight) and using the three-dimensional reconstruction to explain features of the two-dimensional images. The advantages of our proposed method are that it degrades gracefully down to a single image, that it uses linear equations and constraints (allowing the use of convex optimization), that it is amenable to three-dimensional structural biases, and that ambiguity can be expressed precisely (it is possible to “know what we don’t know”). Previously described methods have some, but not all, of these properties.
This paper examines the role of NBTI and PBTI on SRAM Vmin shifts during HTOL stressing and quantifies their impact on reliability lifetime projections in scaled high-k metal gate (HKMG) technologies. Correlation between measured HTOL SRAM Vmin shifts and transistor level parametrics is summarized on both 28nm poly-SiON and HKMG technologies. The paper concludes that the commonly used HTOL acceleration voltage of 1.4xVnom may be excessive in scaled HKMG technologies due to the larger role of PBTI in SRAMs.
This paper presents a recognizer for identifying references to user interface components in online documentation. The recognizer first extracts phrases matching a list of known components, then employs a classifier to reject coincidental matches. We describe why this seemingly straightforward problem is challenging, then show how informal conventions in documentation writing can be leveraged to perform classification. Using the features identified in this paper, our approach achieves an average F1 score of 0.81, and can correctly distinguish between actual command references and coincidental matches in 93.7% of test cases.
This paper introduces query-feature graphs, or QF-graphs. QF-graphs encode associations between high-level descriptions of user goals (articulated as natural language search queries) and the specific features of an interactive system relevant to achieving those goals. For example, a QF-graph for the GIMP graphics manipulation software links the query "GIMP black and white" to the commands "desaturate" and "grayscale." We demonstrate how QF-graphs can be constructed using search query logs, search engine results, web page content, and localization data from interactive systems. An analysis of QF-graphs shows that the associations produced by our approach exhibit levels of accuracy that make them eminently usable in a range of real-world applications. Finally, we present three hypothetical user interface mechanisms that illustrate the potential of QF-graphs: search-driven interaction, dynamic tooltips, and app-to-app analogy search.
People routinely rely on Internet search engines to support their use of interactive systems: they issue queries to learn how to accomplish tasks, troubleshoot problems, and otherwise educate themselves on products. Given this common behavior, we argue that search query logs can usefully augment traditional usability methods by revealing the primary tasks and needs of a product's user population. We term this use of search query logs CUTS - characterizing usability through search. In this paper, we introduce CUTS and describe an automated process for harvesting, ordering, labeling, filtering, and grouping search queries related to a given product. Importantly, this data set can be assembled in minutes, is timely, has a high degree of ecological validity, and is arguably less prone to self-selection bias than data gathered via traditional usability methods. We demonstrate the utility of this approach by applying it to a number of popular software and hardware systems.
Recognizing hand drawn mathematical matrices on tablet computers has proven to be a particularly challenging task. While individual expression recognition can be simplified by assuming the entire content is a single semantic construct, a single math expression, a matrix is composed of multiple expressions arranged in rows and columns. These expressions must first be segmented into matrix elements, and then each individual matrix element expression must be recognized. In this work, we show how a simple algorithm on in-air (i.e. non-inking) strokes can be used to analyze the drawing order of a matrix. Once the drawing order is recognized, we show how outlier analysis on in-air packets gives rapid, reliable segmentation of matrix elements.
Driven by the increasing availability of low-cost sensing hardware, gesture-based input is quickly becoming a viable form of interaction for a variety of applications. Electronic presentations (e.g., PowerPoint, Keynote) have long been seen as a natural fit for this form of interaction. However, despite 20 years of prototyping such systems, little is known about how gesture-based input affects presentation dynamics, or how it can be best applied in this context. Instead, past work has focused almost exclusively on recognition algorithms. This paper explicitly addresses these gaps in the literature. Through observations of real-world practices, we first describe the types of gestures presenters naturally make and the purposes these gestures serve when presenting content. We then introduce Maestro, a gesture-based presentation system explicitly designed to support and enhance these existing practices. Finally, we describe the results of a real-world field study in which Maestro was evaluated in a classroom setting for several weeks. Our results indicate that gestures which enable direct interaction with slide content are the most natural fit for this input modality. In contrast, we found that using gestures to navigate slides (the most common implementation in all prior systems) has significant drawbacks. Our results also show how gesture-based input can noticeably alter presentation dynamics, often in ways that are not desirable.