The H-1B program authorizes non-immigrant visas under which skilled foreign workers may be employed in the U.S., typically in computer-related positions. Congress greatly expanded the program in 1998 and then again in 2000, in response to heavy pressure from industry, which claimed a desperate software labor shortage. After presenting an overview of the H-1B program in Parts II and III, the Article will show in Part IV that these shortage claims are not supported by the data. Part V will then show that the industry's motivation for hiring H-lBs is primarily a desire for cheap, compliant labor. The Article then discusses the adverse impacts of the H-1B program on various segments of the American computer-related labor force in Part VI, and presents proposals for reforms in Part VII.
The public page is a popular online social community platform. These pages form a network by liking each other. Location classification of public pages has been studied at the country and state levels. In this paper, we explored the task of public page classification by cities within California. We introduced a virtual geographic structure for city clusters resembling counties in California. We developed a clustering algorithm that leverages the confusion matrix from flat city classification to construct the virtual geographic city structure. Then, adopting a two-stage hierarchical classification strategy-first classifying pages by city cluster and then within clusters by city-we enhanced the accuracy from 0.6928 of flat city classification to 0.8014.
A novel variation of the data swapping approach to statistical disclosure control is presented, aimed particularly at preservation of multivariate relations in the original dataset. A theorem is proved in support of the method, and extensive empirical investigation is reported.
The old empathetic adage, ``Walk a mile in their shoes,'' asks that one imagine the difficulties others may face. This suggests a new ML counterfactual fairness criterion, based on a \textit{group} level: How would members of a nonprotected group fare if their group were subject to conditions in some protected group? Instead of asking what sentence would a particular Caucasian convict receive if he were Black, take that notion to entire groups; e.g. how would the average sentence for all White convicts change if they were Black, but with their same White characteristics, e.g. same number of prior convictions? We frame the problem and study it empirically, for different datasets. Our approach also is a solution to the problem of covariate correlation with sensitive attributes.
A number of methods have been introduced for the fair ML issue, most of them complex and many of them very specific to the underlying ML moethodology. Here we introduce a new approach that is simple, easily explained, and potentially applicable to a number of standard ML algorithms. Explicitly Deweighted Features (EDF) reduces the impact of each feature among the proxies of sensitive variables, allowing a different amount of deweighting applied to each such feature. The user specifies the deweighting hyperparameters, to achieve a given point in the Utility/Fairness tradeoff spectrum. We also introduce a new, simple criterion for evaluating the degree of protection afforded by any fair ML method.
The k‐nearest neighbors (k‐NN) method is one of the oldest statistical/machine learning techniques. It is included in virtually every major package, such as caret, parsnip, mlr3 and scikit‐learn. Yet those packages do not go beyond the basics. With today's high‐speed computation capability, k‐NN can be made much more powerful. Here, we present directions in which that can be done.
With technological advances leading to an increase in mechanisms for image tampering, fraud detection methods must continue to be upgraded to match their sophistication. One problem with current methods is that they require prior knowledge of the method of forgery in order to determine which features to extract from the image to localize the region of interest. When a machine learning algorithm is used to learn different types of tampering from a large set of various image types, with a large enough database we can easily classify which images are tampered (by training on the entire image feature map for each image) [28]. However, we still are left with the question of which features to train on, and how to localize the manipulation. To solve this, object detection networks such as Faster R-CNN [22], which combine an RPN (Region Proposal Network) with a CNN, have recently been adapted to fraud detection by utilizing their ability to propose bounding boxes for objects of interest to localize the tampering artifacts [29]. By making use of the computational powers of todays GPUs this method also achieves a fast run-time and higher accuracy than the top current methods such as noise analysis, ELA (Error Level Analysis), or CFA (Color Filter Array) [29]. In this work, a multi-linear Faster RCNN network will be applied similarly but with the second stream having an input of the ELA JPEG compression level mask. This is shown to provide even higher accuracy by adding training features from the segmented image map to the network.
Today the terms machine learning (ML) and Big Data are closely correlated. This, and the complexity of many ML algorithms, motivates a search for fast parallel computation methods. A further motivating factor is a need to deal with memory size limitations, especially for the moderately-sized machines common in many ML applications. In addition, it is desirable to develop generally applicable methods, rather than needing to develop a different parallel approach for every ML algorithm. In this work, we apply a technique we call Software Alchemy to ML. We are particularly interested in ML for recommender systems, and explore the feasibility of SA in that context.
Despite the success of neural networks (NNs), there is still a concern among many over their "black box" nature. Why do they work? Here we present a simple analytic argument that NNs are in fact essentially polynomial regression models. This view will have various implications for NNs, e.g. providing an explanation for why convergence problems arise in NNs, and it gives rough guidance on avoiding overfitting. In addition, we use this phenomenon to predict and confirm a multicollinearity property of NNs not previously reported in the literature. Most importantly, given this loose correspondence, one may choose to routinely use polynomial models instead of NNs, thus avoiding some major problems of the latter, such as having to set many tuning parameters and dealing with convergence issues. We present a number of empirical results; in each case, the accuracy of the polynomial approach matches or exceeds that of NN approaches. A many-featured, open-source software package, polyreg, is available.
Two electro-optical computer interface embodiments provide for one-way read only and two-way optical read and write. The two-way embodiment includes a main module having a shared memory, and a proces sor/controller and a main bus. A plurality of local pro cessor modules each includes a local memory, a local processor, and a local bus. The processors are electri cally joined by control conductors which provide for coordination and timing between local processors and the main processor. Each memory array has a film de posited on it by the Langmuir/Blodgett technique. The memory arrays are each illuminated by a pulsed laser or Q-switched laser. The film is responsive to the electric fields in the memory array cells for modulating the illumination light. The image is then read onto other memory arrays which are responsive to the illumination for transferring the data between memories.
In recent years there has been widespread concern in the scientific community over a reproducibility crisis. Among the major causes that have been identified is statistical: In many scientific research the statistical analysis (including data preparation) suffers from a lack of transparency and methodological problems, major obstructions to reproducibility. The revisit package aims toward remedying this problem, by generating a "software paper trail" of the statistical operations applied to a dataset. This record can be "replayed" for verification purposes, as well as be modified to enable alternative analyses. The software also issues warnings of certain kinds of potential errors in statistical methodology, again related to the reproducibility issue.
Karl Levitt合作论文数Department of Computer Science University of California Davis1
Meng Chang Chen合作论文数Institute of Information Science;Academia Sinica1