Publications
2026
- TOSEMJournalJiangnan Huang, and Bin LinACM Trans. Softw. Eng. Methodol. Sep 2026
GitHub Actions, a built-in CI/CD solution of GitHub, is increasingly popular among developers for automating software development workflows. It has been observed that when automated workflow execution fails, developers sometimes only rerun the workflow or failed job without any modifications to the repository. The widespread use of these reruns has consumed considerable computing resources and raised concerns regarding the reliability and consistency of these workflows. Understanding how developers rerun GitHub Actions workflows and the rationale behind the rerun can provide valuable insights to further improve the reliability and efficiency of the software development process.In this work, we conducted an empirical study on 3,320 open source Java repositories to understand how developers rerun GitHub Actions workflows and quantify both wasted time and computing resources. We further studied the cases where workflow reruns lead to successful outcomes and manually analyzed the reasons behind the workflow execution flakiness. Based on our findings, we tested four machine learning–based models to predict the workflow execution outcome, aiming to reduce the potential resource waste caused by workflow reruns. Our study presents how developers deal with GitHub Actions execution failures, pinpoints root causes of workflow execution flakiness, and offers actionable insights to improve CI/CD workflow reliability and efficiency.
@article{10.1145/3795771, author = {Huang, Jiangnan and Lin, Bin}, title = {On the Reruns of GitHub Actions Workflows}, year = {2026}, issue_date = {October 2026}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, volume = {35}, number = {10}, issn = {1049-331X}, url = {https://doi.org/10.1145/3795771}, doi = {10.1145/3795771}, journal = {ACM Trans. Softw. Eng. Methodol.}, month = sep, articleno = {328}, numpages = {32}, keywords = {GitHub Actions, CI/CD, Software Repositories}, } - ICSE2026ConferenceJiangnan HuangIn Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering Sep 2026
CI/CD pipelines form a core component of contemporary software engineering, automating key stages of the development lifecycle. Since its release in 2019, GitHub Actions (GHA) has rapidly emerged as the dominant CI/CD platform on GitHub, automating millions of workflows daily. However, as the ecosystem grows, developers face persistent challenges in automating GHA workflow creation, mitigating workflow failures, and ensuring security. To address these challenges, my doctoral research systematically examines how developers construct, maintain, and secure GHA workflows in practice, and proposes learning-based approaches to improve their automation quality. The research includes two empirical studies on workflow reliability (reruns and flakiness) and security (evolution of developers’ practices), along with two learning-based frameworks: CIGAR for recommending reusable Actions and FLOWSTEP for generating step-aware workflows. Together, they aim to build an intelligent and secure automation framework for GitHub Actions integrating reliability, security, and semantic understanding.
@inproceedings{10.1145/3774748.3787630, author = {Huang, Jiangnan}, title = {Toward Intelligent, Reliable, and Secure Automation of CI/CD Workflows}, year = {2026}, isbn = {9798400722967}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3774748.3787630}, doi = {10.1145/3774748.3787630}, booktitle = {Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering}, pages = {150–152}, numpages = {3}, keywords = {GitHub Actions, CI/CD, Automation, Security, Reliability}, location = {Rio de Janeiro, Brazil}, series = {ICSE-Companion '26}, } - TOSEMJournalLianyu Zheng, Shuang Li, Xi Huang, and 4 more authorsACM Trans. Softw. Eng. Methodol. Apr 2026
GitHub actions (GHA), a built-in continuous integration and continuous delivery (CI/CD) service of GitHub, has been widely adopted by developers, streamlining the automation of software development workflows. Despite its popularity, failures frequently occur during GHA workflow executions. Fixing these failures often requires significant human effort, and unsuccessful workflow executions waste computing resources. Understanding the reasons behind workflow failures could provide valuable insights for troubleshooting the existing issues of CI/CD and further improving the development process.In this article, we present an empirical study to reveal the reasons behind GHA workflow failures. By manually analyzing 375 failed workflow executions across 260 open-source Java projects, we built a comprehensive taxonomy categorizing the common failure types. The taxonomy was further validated by surveying 151 developers. This study is the first empirical work to analyze GHA workflow failures, bringing valuable knowledge to the field of continuous integration in software engineering. Moreover, our taxonomy and survey results not only underscore the critical need for better tools and practices to mitigate these failures but also indicate the directions to enhance the efficiency and reliability of CI/CD pipelines.
@article{10.1145/3749371, author = {Zheng, Lianyu and Li, Shuang and Huang, Xi and Huang, Jiangnan and Lin, Bin and Chen, Jinfu and Xuan, Jifeng}, title = {Why Do GitHub Actions Workflows Fail? An Empirical Study}, year = {2026}, issue_date = {May 2026}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, volume = {35}, number = {5}, issn = {1049-331X}, url = {https://doi.org/10.1145/3749371}, doi = {10.1145/3749371}, journal = {ACM Trans. Softw. Eng. Methodol.}, month = apr, articleno = {139}, numpages = {29}, keywords = {GitHub Workflow, Continuous Integration}, }
2025
- ICPC2025ConferenceJiangnan Huang, and Bin LinIn Proceedings of the 33rd International Conference on Program Comprehension Apr 2025
GitHub Actions, a built-in CI/CD service of GitHub released in 2019, has become one of the most widely adopted tools among developers for automating software development workflows. This popularity, however, brings security challenges, as vulnerable workflows can expose repositories and the software supply chains to significant risks. Existing studies have highlighted several types of potential security issues. Over the past few years, GitHub has been constantly promoting better security practices and developers have gained experience in using GitHub Actions. Investigating how developers’ practices for handling GitHub Actions security have changed over time could offer valuable insights for further strengthening the security of these workflows. In this study, we analyzed non-optimal security practices in 18,938 workflows from 5,246 active GitHub repositories. By comparing the prevalence of issues spotted in two different years (2022 and 2024), we find that some undesired practices still widely exist in repositories. However, some progress has been observed, such as the significant reduce of permission misconfigurations.
@inproceedings{Huang2025a, author = {Huang, Jiangnan and Lin, Bin}, title = {Revisiting Security Practices for GitHub Actions Workflows}, booktitle = {Proceedings of the 33rd International Conference on Program Comprehension}, series = {ICPC '25}, year = {2025}, location = {Ottawa, Canada}, pages = {accepted}, numpages = {5}, }
2023
- SCAM2023ConferenceCIGAR: Contrastive Learning for GitHub Action Recommendation Best Paper AwardJiangnan Huang, and Bin LinSource Code Analysis and Manipulation Apr 2023
GitHub Actions was introduced in 2019 as an integrated solution for CI/CD to automate software development workflow. Since then, it has gained tremendous popularity among developers. In a GitHub Actions workflow, actions refer to custom applications for performing complex but frequently repeated tasks. Actions can be typically found in GitHub Marketplace or public GitHub repositories. Prior studies have already disclosed that developers often reuse actions to reduce double work and improve productivity. However, it is not trivial for developers, especially novices, to figure out which action to reuse due to the large number of actions available and the limited search functionality GitHub Marketplace provides. To address this issue, we propose CIGAR (ContrastIve learning for GitHub Action Recommendation). Given the textual description of a task developers want to execute, CIGAR will recommend the most relevant actions. CIGAR exploits a pre-trained RoBERTa model to convert sequences of words into high-dimensional vector representations, and is fine tuned through a contrastive learning objective. The performance of CIGAR was evaluated on a novel dataset curated based on prior research, and the results demonstrate that CIGAR can reliably recommend actions needed by developers and significantly outperforms the GitHub Marketplace search engine. Our study indicates the promise of employing contrastive learning for GitHub action recommendation. The promising performance achieved can potentially drive a wider adoption of GitHub Actions and facilitate the automation of software development workflows.
@article{a89159487haha, journal = {Source Code Analysis and Manipulation}, year = {2023}, author = {Huang, Jiangnan and Lin, Bin}, title = {CIGAR: Contrastive Learning for GitHub Action Recommendation}, award = {Best Paper Award}, }
2022
- AlgorithmsJournalJoseph Pedersen, Rafael Muñoz-Gómez, Jiangnan Huang, and 3 more authorsAlgorithms Apr 2022
We address the problem of defending predictive models, such as machine learning classifiers (Defender models), against membership inference attacks, in both the black-box and white-box setting, when the trainer and the trained model are publicly released. The Defender aims at optimizing a dual objective: utility and privacy. Privacy is evaluated with the membership prediction error of a so-called “Leave-Two-Unlabeled” LTU Attacker, having access to all of the Defender and Reserved data, except for the membership label of one sample from each, giving the strongest possible attack scenario. We prove that, under certain conditions, even a “naïve” LTU Attacker can achieve lower bounds on privacy loss with simple attack strategies, leading to concrete necessary conditions to protect privacy, including: preventing over-fitting and adding some amount of randomness. This attack is straightforward to implement against any model trainer, and we demonstrate its performance against MemGaurd. However, we also show that such a naïve LTU Attacker can fail to attack the privacy of models known to be vulnerable in the literature, demonstrating that knowledge must be complemented with strong attack strategies to turn the LTU Attacker into a powerful means of evaluating privacy. The LTU Attacker can incorporate any existing attack strategy to compute individual privacy scores for each training sample. Our experiments on the QMNIST, CIFAR-10, and Location-30 datasets validate our theoretical results and confirm the roles of over-fitting prevention and randomness in the algorithms to protect against privacy attacks.
@article{a15070254, author = {Pedersen, Joseph and Muñoz-Gómez, Rafael and Huang, Jiangnan and Sun, Haozhe and Tu, Wei-Wei and Guyon, Isabelle}, title = {LTU Attacker for Membership Inference}, journal = {Algorithms}, volume = {15}, year = {2022}, number = {7}, article-number = {254}, url = {https://www.mdpi.com/1999-4893/15/7/254}, issn = {1999-4893}, doi = {10.3390/a15070254}, }
2021
- OLA2021ConferenceJiangnan Huang, Zixi Chen, and Nicolas DupinIn Optimization and Learning Apr 2021
Having N points in a planar Pareto Front (2D PF), k-means and k-medoids are solvable in O(N^3) time by dynamic programming algorithms. Standard local search approaches, PAM and Lloyd’s heuristics, are investigated in the 2D PF case to solve faster large instances. Specific initialization strategies related to 2D PF cases are implemented with the generic ones (Forgy’s, Hartigans, k-means++). Applying PAM and Lloyd’s local search iterations, the quality of local minimums are compared with optimal values. Numerical results are computed using generated instances, which were made public. This study highlights that local minimums of a poor quality exist for 2D PF cases. A parallel or multi-start heuristic using four initialization strategies improves the accuracy to avoid poor local optimums. Perspectives are still open to improve local search heuristics for the specific 2D PF cases.
@inproceedings{10.1007/978-3-030-85672-4_2, author = {Huang, Jiangnan and Chen, Zixi and Dupin, Nicolas}, title = {Comparing Local Search Initialization for K-Means and K-Medoids Clustering in a Planar Pareto Front, a Computational Study}, booktitle = {Optimization and Learning}, year = {2021}, publisher = {Springer International Publishing}, address = {Cham}, pages = {14--28}, isbn = {978-3-030-85672-4}, }