
Malicious JavaScript detection using machine learning models has shown many great results over the years. However, real-world data only has a small fraction of malicious JavaScript, making it an imbalanced dataset. Many of the previous techniques ignore most of the benign samples and focus on training a machine learning model with a balanced dataset. This paper proposes a Doc2Vec-based filter model that can quickly classify JavaScript malware using Natural Language Processing (NLP) and feature re-sampling. The feature of the JavaScript file will be converted into vector form and used to train the classifiers. Doc2Vec, a NLP model used for documents is used to create feature vectors from the datasets. In this paper, the total features of the benign samples will be reduced using a combination of word vector and clustering model. Random seed oversampling will be used to generate new training malicious data based on the original training dataset. We evaluate our models with a dataset of over 30,000 samples obtained from top popular websites, PhishTank, and GitHub. The experimental result shows that the best f1-score achieves at 0.99 with the MLP classifier.
Recently, it was shown that the Chudnovsky-type algorithms over the projective line enable a generic recursive construction of multiplication algorithms in finite fields $$\mathbb {F}_{q^n}$$ . This construction is based on the method introduced by D. V. and G. V. Chudnovsky (1987), that generalizes polynomial interpolation to the use of algebraic curves, and specialized to the case of the projective line. More precisely, using evaluation at places of increasing degrees of the rational function field over $$\mathbb {F}_q$$ , it has been proven that one can construct in polynomial time algorithms with a quasi-linear bilinear complexity with respect to the extension degree n. In this paper, we discuss how the use of evaluation with multiplicity, using generalized evaluation maps, can improve this construction.
Research on codes over finite rings has intensified after the discovery that some of the best binary nonlinear codes can be obtained as images of ℤ_4 -linear codes. Codes over various finite rings have been a subject of much research in coding theory after this discovery. Many of these rings are extensions of ℤ_4 and numerous new linear codes over ℤ_4 have been found in the last decade. Due to the special place of ℤ_4 , an online database of ℤ_4 codes was created in 2007. The original database on ℤ_4 codes recently became unavailable. The purpose of this paper is to introduce a new and updated database of ℤ_4 codes. We have made major updates to the original database by adding 8699 new linear codes over ℤ_4 . These codes have been found through exhaustive computer searches on cyclic codes and by an implementation of the ASR search algorithm that has been remarkably fruitful to obtain new linear codes from the class of quasi-cyclic (QC) and quasi-twisted (QT) codes over finite fields. We made modifications to the ASR algorithm to make it work over ℤ_4 . The initial database contained few codes that were not free. We have now added a large number of non-free codes. In fact, of the 8699 codes we have added, 7631 of them are non-free. Our database will be further updated by incorporating data from [16]. We also state an open problem of great theoretical interest that arose from computational observations.
As the development of quantum machines is booming and would threaten our standard cryptography algorithms, a transition period is necessary for the protection of the data processed by our classical machines as well before the arrival of theses machines as after. Recently, to get ahead of the curve, the National Institute of Standards and Technology (NIST) launched the Post Quantum Cryptography Standardization Project, started since late 2016. Among finalists, 3 promising code-theoretic finalist candidates, Classic McEliece, BIKE, and HQC are sent to the fourth round. In this work, to reduce classical McEliece key size without loss of security, we present a new key generation algorithm by introducing new family of codes called quasi-centrosymmetric Goppa codes with a moderate key size for storage optimisation. We also have characterized these codes in the case where the parity matrix is in Cauchy form by giving an algorithm to build them. We ended up giving a detailed analysis of the security against the most known structural attacks by giving the new complexities.
QCB is a proposal for a post-quantum secure, rate-one authenticated encryption with associated data scheme (AEAD) based on classical OCB3 and CB, which are vulnerable against a quantum adversary in the Q2 setting. The authors of QCB prove integrity under plus-one unforgeability, whereas the proof of the stronger definition of blind unforgeability has been left as an open problem. After a short overview of QCB and the current state of security definitions for authentication, this work proves blind unforgeability of QCB. Finally, the strategy of using tweakable block ciphers in authenticated encryption is generalised to a generic blindly unforgeable AEAD model.
The MPC-in-the-head introduced in [ IKOS07 ] has established itself as an important paradigm to design efficient digital signatures. For instance, it has been leveraged in the Picnic scheme [ CDG+20 ] that reached the third round of the NIST Post-Quantum Cryptography Standardization process. In addition, it has been used in [ Beu20 ] to introduce the Proof of Knowledge (PoK) with Helper paradigm. This construction permits to design shorter signatures but induces a non negligible performance overhead as it uses cut-and-choose. In this paper, we introduce the PoK leveraging structure paradigm along with its associated challenge space amplification technique. Our new approach to design PoK brings some improvements over the PoK with Helper one. Indeed, we show how one can substitute the Helper in these constructions by leveraging the underlying structure of the considered problem. This new approach does not suffer from the performance overhead inherent to the PoK with Helper paradigm hence offers different trade-offs between security, signature sizes and performances. In addition, we also present four new post-quantum signature schemes. The first one is based on a new PoK with Helper for the Syndrome Decoding problem. It relies on ideas from [ BGKM22 ] and [ FJR21 ] and improve the latter using a new technique that can be seen as performing some cut-and-choose with a meet in the middle approach. The three other signatures are based on our new PoK leveraging structure approach and as such illustrate its versatility. Indeed, we provide new PoK related to the Permuted Kernel Problem (), Syndrome Decoding () problem and Rank Syndrome Decoding problem. Considering (public key + signature), we get sizes below 9 kB for our signature related to the problem, below 15 kB for our signature related to the problem and below 7 kB for our signature related to the problem. These new constructions are particularly interesting presently as the NIST has recently announced its plan to reopen the signature track of its Post-Quantum Cryptography Standardization process.
The th order Gowers norm of a Boolean function is a measure of its resistance to r th order approximations. Gowers, Green, and Tao presented the connection between Gowers uniformity norm of a function and its correlation with polynomials of a certain degree. Gowers and norms measure the resistance of a Boolean function against linear and quadratic approximations, respectively. This paper presents computational results on the Gowers and norms of some known 4, 5, and 6-bit S-Boxes used in the present-day block ciphers. It is observed that there are S-Boxes having the same algebraic degree, differential uniformity, and linearity, but their Gowers norms values are different. Equivalently, they possess different strengths against linear and quadratic cryptanalysis.
Most cryptographers believe our modern systems and proven secure protocols cannot be broken. Some are now convincing Treasure Departments to make their own versions of Bitcoin. Today cryptosystems are considered secure as long as academics have not broken them. Earlier we argued that this approach might be badly flawed, by presenting a minority viewpoint. In this paper, we look at what the history of Al-Andalusia might teach us in regard to our research community. We also wonder how “open” our so called “open” research is.
We introduce the scheme of a keyed hash function, also called message authentication code (MAC) based on previously given latin squares and binary linear error-correcting codes. We investigate the properties of the introduced scheme regarding its security and applicability. Further, we give a possible application in a smart home environment, especially for opening the entrance door to a house by using a mobile device.
We consider a non-interactive secure computation protocol that we call multi-input non-interactive functional encryption (MINI-FE). In a MINI-FE protocol for some class of functionalities F , N users independently and non-interactively setup their own pairs of public- and secret-keys. Each user i∈{1, … ,N} , with knowledge of the public-keys of all other participants, its own secret-key, an identifier and a functionality F, can encode its own input x_i to produce a ciphertext _i . There is a public evaluation function that, given N ciphertexts _1,… ,_N for the same identifier and functionality F, outputs F(x_1,… ,x_n) . Moreover, the same public keys can be reused for an unbounded number of computations (for different identifiers). The security essentially guarantees that for any two tuples of inputs that map to the same value under a functionality, the corresponding tuples of ciphertexts are computationally indistinguishable to adversaries that can also corrupt a subset of the users. MINI-FE shares some similarities with decentralized multi-client functional encryption (DMCFE) [Chotard et al. - Asiacrypt ’18] but, unlike DMCFE and alike dynamic DMCFE (DDMCFE) [Chotard et al. - Crypto ’20], in MINI-FE there is no interaction in the setup phase. Unlike DDMCFE and similarly to traditional secure computation, there is no concept of token and thus the leakage of information is limited to the execution of each computation for a given identifier. In MINI-FE each user can completely work independently and the entire protocol can be executed over a broadcast channel (e.g., a distributed ledger) with just a single message from each user in both the setup and encoding phases. We give an instantiation of a MINI-FE protocol for the Inner-Product functionality from bilinear groups; previous constructions were only known for the summation functionality. We show applications of our protocol to private stream aggregation and secure quadratic voting.
We verify the achievement of information-theoretic security of Visual Cryptography (VC) based on the detailed attack scenario. In addition, practical VCs use pseudo-random permutation (PRP) as a random shuffle, which we also verify in this case. As a result, practical VCs only have computational security and require true random number generators (TRNGs) to achieve information-theoretic security. We derive a rational evaluation standard for VC’s security from the detailed attack scenario. Based on it, we consider an attack method using multiple shared images that is more effective than the brute-force search. In this study, we execute computer simulations of XOR differential attack on two shared images and observe the bias on the appearance frequency of the differential output. From these above results, we can conclude that it is necessary to generate the basis matrices with the same Hamming weight for all rows to guarantee the security of VCs.
This paper proposes two methods that combine high error correcting capability with security enhancement to enable cryptographic communication even under high noise. The first method is a combination of symmetric key cryptography and Shortened LDPC, which enables two-way communication. It can be regarded as one type of mode of operation. The second method combines the McEliece method and Shortened QC-MDPC to realize one-way communication. It has the advantage of fast processing speeds compared to general asymmetric key cryptography and the ability to centrally manage key updates for many IoT modules. We performed computer simulations and analysed practical parameterization and security enhancement. Both methods are found to provide sufficient security and are expected to have a wide range of applications.
Technology has always had a significant role in medicine. Modern advancements in technology has revolutionized the health industry and totally changed the way we currently undertake healthcare. Telecare medicine information system (TMIS) is among the best new technological innovations shaping the medical field and taking its digitization to another logical step. It has made distance redundant and enabled patients to consult with specialists practically anywhere on the globe, which can be a lifesaver in emergency situations where immediate care is required. However, within TMISs, personal and sensitive information about patients and health professionals are exchanged over public networks and may be subject to several security attacks. Therefore, a wide variety of schemes have been proposed in literature to deal with user authentication and key agreement issues, yet, the majority of these works fail to achieve the essential security features. In this paper, we first provide a review of Ostad-Sharif et al. scheme and demonstrate its vulnerability to key compromise password and key compromise impersonation attacks. Consequently, to cover these security weaknesses, a robust mutual and anonymous three-factor authentication scheme with key agreement protocol is proposed for lightweight application in TMISs. Proof of security is given through an informal security analysis and a formal security verification using AVISPA. Our performance evaluation reveals that the proposed scheme is cost efficient in terms of computation, communication, storage and network impact overheads, as compared to other related recent schemes.
4-adic complexity is an important characteristic of the unpredictability of a sequence. It is defined as the smallest order of feedback with carry shift register that can generate the whole sequence. In this paper, we estimate the symmetric 4-adic complexity of two classes of quaternary sequences with period $$2p^n$$ . These sequences are constructed on generalized cyclotomic classes of order four and have a high linear complexity over the finite ring and over the field of order four. We show that the 4-adic complexity is good enough to resist the attack of the rational approximation algorithm.
Meneghetti, Picozzi, and Tognolini proposed a Hamming-metric code-based digital signature scheme from QC-LDPC codes in [ 10 ]. Using quasi cyclic codes enables the scheme to have small key sizes, while the structure of LDPC codes helps to achieve good overall performance. In this paper, we discuss the security of this code-based signature scheme. In particular, we present a partial key recovery attack that uses a statistical method to recover an important part of the secret key. Using a proof-of-concept Sagemath implementation, we are able to recover part of the secret key using as few as 25 signatures in less than 2 minutes for all proposed parameter sets. Furthermore, we follow up this partial key recovery attack by a forgery attack.
Substitution Permutation Networks (SPNs) are widely used in the design of modern symmetric cryptographic building blocks. In their Eurocrypt 2016 paper titled ‘Indifferentiability of Confusion-Diffusion Networks’, Dodis et al. theorized such SPNs as Confusion-Diffusion networks and established their provable security in Maurer’s indifferentiability framework. Guo et al. extended this work to non-linear Confusion-Diffusion networks (NLCDNs), i.e., networks using non-linear permutation layers, in weaker indifferentiability settings. The authors provided a security proof in the sequential indifferentiability model for the 3-round NLCDN and exhibited the tightness of the positive result by providing an (incorrect) attack on the 2-round NLCDN. In this paper, we provide a corrected attack on the 2-round NLCDN. Our attack on the 2-round CDN is primitive-construction-sequential, implying that the construction is not secure even in the weaker sequential indifferentiability setting of Mandal et al. In their paper titled ‘Revisiting Cascade Ciphers in Indifferentiability Setting’, Guo et al. showed that four stages are necessary and sufficient to realize an ideal $$(2\kappa ,n)$$ -block cipher using the cascade of independent ideal $$(\kappa ,n)$$ -block ciphers with two alternated independent keys, in the indifferentiability paradigm (where a (k, n)-blockcipher has k-bit key space and n-bit message space). As part of their negative results, Guo et al. provided attacks for the 2-round and 3-round cascade constructions with two alternating keys. Further, they gave a heuristic outline of an attack on the 3-round cascade construction with (certain) stronger key schedules. As the second half of this paper, we formalize the attack explored by Guo et al. on the 3-round cascade construction with stronger key schedules and extend the same to any 2n-bit to 3n-bit non-idealized key scheduling function.
An accumulator is a cryptographic protocol that compresses a set of inputs into a short string of a certain size and can efficiently prove that the compressed set contains a particular input element. Accumulators have been actively studied in recent years and are used to streamline various protocols such as membership rosters, zero-knowledge proofs, group signatures, and blockchains. Libert et al. proposed a Merkle tree-based accumulator using lattice cryptography, one of the post-quantum cryptography. They proposed an accumulator with logarithmic time complexity for the verification algorithm. Ling et al. proposed an accumulator that satisfies logarithmic time updating lists. However, no algorithm has been proposed thus far that satisfies constant time updating lists and constant time verification based on the lattice-based accumulator. In this study, we propose an accumulator based on lattice that satisfies constant-time verification and constant-time updating lists for the first time. In our proposed accumulator, the bit length of the witness associated with each element is independent of the number of elements in the list. We developed techniques that use the Partial Fourier Recovery problem instead of the Merkle tree. We also prove that the proposed accumulator satisfies the security requirements of an accumulator scheme. Finally, to demonstrate that our proposed accumulator is more practical, we compared it with other lattice-based accumulators. The proposed accumulator scheme can be incorporated into membership list management, zero-knowledge proof, group signature, and blockchain to realize more efficient applications.
Quantum computers are a threat to the current standards for secure communication. The Datagram Transport Layer Security (DTLS) protocol is a common protocol used by Internet of Things (IoT) devices that will be broken by such computers. Although quantum computers are yet to become commercially available, IoT devices are generally long-lived. Thus the transition to quantum secure cryptography, as soon as possible, is necessary. IoT devices are generally resource-constrained and Post-Quantum (PQ) cryptography is often more resource intensive computationally compared to current cryptographic standards, adding to the complexity of the transition. In this paper, we propose a PQ version of DTLS 1.3 in IoT, at some additional costs. We first identify a suitable PQ digital signature scheme and Key Encapsulation Mechanism (KEM) to be used in a PQ version of the DTLS protocol. Using the selected PQ algorithms, we implement and evaluate a full PQ DTLS 1.3 handshake on a Raspberry Pi 4B. We find that CPU usage is actually lower compared to current cryptographic schemes used in DTLS 1.3. We notice a significant increase of up to 6x as many packets sent when establishing a connection, depending on the security level. Moreover, memory usage is significantly greater, requiring at least an extra 800 KiB of memory to connect 100 devices.
Click fraud, the manipulation of online advertisement traffic figures, is becoming a major concern for businesses that advertise online. This can lead to financial losses and inaccurate click statistics. To address this problem, it is essential to have a reliable method for identifying click fraud. This includes distinguishing between legitimate clicks made by users and fraudulent clicks generated by bots or other software, which enables companies to advertise their products safely. The XGBoost model was trained on the TalkingData AdTracking Fraud Detection dataset from Kaggle, using binary classification to predict the likelihood of a click being fraudulent. The importance of data treatment was taken into consideration in the model training process, by carefully preprocessing and cleaning the data before feeding it into the model. This helped to improve the accuracy and performance of the model by reaching an AUC of 0.96 and LogLoss of 0.15.
At CHES'2021, a chosen ciphertext attack combined with belief propagation which can recover the long-term secret key of CRYSTALS-Kyber from side-channel information of the number theoretic transform (NTT) computations was presented. The attack requires k traces from the inverse NTT step of decryption, where k is the module rank, for a noise tolerance $$\sigma \le 1.2$$ in the Hamming weight (HW) leakage on simulated data. In this paper, we present an attack which can recover the secret key of CRYSTALS-Kyber from k chosen ciphertexts using side-channel information of the Barret reduction and message decoding steps of decryption, for $$k \in \{3,4\}$$ . The key novel idea is to create a unique mapping between the secret key coefficients and multiple intermediate variables of these procedures. The redundancy in the mapping patterns enables us to detect errors in the secret key coefficients recovered from side-channel information. We demonstrate the attack on the example of a software implementation of Kyber-768 in ARM Cortex-M4 CPU using deep learning-based power analysis.