您的瀏覽器不支援JavaScript語法,網站的部份功能在JavaScript沒有啟用的狀態下無法正常使用。

Institute of Information Science, Academia Sinica

Research

Print

Press Ctrl+P to print from browser

Recent Research Results

:::

Search-Tree Scaling in Parallel Monte Carlo Tree Search

Annual Conference on Neural Information Processing Systems (NeurIPS), December 2026

Scott Cheng, Meng-Yu Tsai, Ding-Yong Hong, Mahmut Kandemir

Scott Cheng Ding-Yong Hong

Abstract

The exploration-exploitation tradeoff is a fundamental challenge in reinforcement learning. In Monte Carlo Tree Search (MCTS), this tradeoff is balanced explicitly by the pUCT exploration coefficient and implicitly by parallelism when the search is combined with neural network evaluations. However, the exploration coefficient is typically tuned as an algorithmic hyperparameter, while the degree of parallelism is often treated as a system detail. To this end, we characterize how these two factors shape the search-tree structure and jointly affect performance scaling. We prove that, under regularity conditions, the depth of the optimal path grows as $\Theta(\sqrt{n}/\log n)$ for $n$ simulations under the MuZero exploration coefficient formula. For tree-parallel MCTS, we prove that the search depth preserves the same asymptotic order when $T=o(\sqrt{n})$, while sufficiently large parallelism introduces contention and reduces the depth growth rate to $\Theta(n/T)$ for $T$ threads. Our empirical results indicate that, in the Go experiment, each doubling of the thread count reduces the best-performing exploration coefficient by approximately 0.19. Overall, our characterization connects the search-tree structure with the exploration-exploitation behavior in pUCT, providing insight into how exploration coefficients and parallelism can be jointly tuned.

CARE: Joint Compression with Adaptive Recovery for Large Language Models

The IEEE International Conference on Parallel and Distributed Systems (ICPADS), November 2026

Fen-Yu Hsieh, Ding-Yong Hong, Jan-Jan Wu

Fen-Yu Hsieh Ding-Yong Hong Jan-Jan Wu

Abstract

Large Language Models (LLMs) have achieved remarkable success across a wide range of complex language tasks. However, deploying these models on resource-constrained platforms remains highly inefficient due to their massive parameter sizes. Although model compression techniques such as pruning and quantization have individually proven effective in reducing model size and inference cost, combining them into a unified method often leads to significant performance degradation, because pruning and quantization impose fundamentally conflicting constraints on the underlying weight distribution. To address this challenge, we propose {\em CARE}, a joint pruning and quantization framework that integrates low-bit weight quantization with semi-structured sparsity while maintaining competitive model accuracy. Our approach leverages second-order Hessian information to compensate for errors arising from compression-induced weight perturbations while strictly preserving the sparsity masks. To further mitigate residual compression errors, we incorporate Low-Rank Adaptation (LoRA) modules to recover accuracy after aggressive compression. In addition, we propose a data-driven sensitivity metric that combines weight magnitudes and gradient information to adaptively assign layer-wise LoRA ranks under a fixed parameter budget. Experimental results on LLaMA-2-7B and LLaMA-3-8B demonstrate that our method reduces the memory footprint by 63.1\% over the uncompressed baseline and outperforms the state-of-the-art SLiM method with 4-bit quantization and 2:4 sparsity.

Optimizing Pipeline Parallelism for Deep Learning With Activation Checkpointing

Concurrency and Computation: Practice and Experience, September 2026

Tzu-Hsien Tsai, Ming-Yen Chiang, Ding-Yong Hong, Pangfeng Liu, Jan-Jan Wu

Ding-Yong Hong Jan-Jan Wu

Abstract

This paper accelerates fast-forward model training that applies pipeline parallelism and activation checkpointing implemented by PyTorch. We prove that it is NP-hard to select checkpoints that minimize the training time under a given memory limit and PyTorch resource usage. We design a dynamic programming algorithm to partition a model into stages that is mathematically proven to minimize the maximum training time among stages. We design a dynamic programming algorithm to select checkpoints that are mathematically proven to minimize the stage training time under a memory limit and PyTorch resource usage. We design an FPTAS for checkpoint selection by rounding the training time of layers. Empirically, our algorithms achieve speedups of up to 2.07x over the state-of-the-art algorithms across models of various sizes and structures. The empirical stage training time corroborates that our algorithms achieve a more balanced stage training time across stages, and our checkpoints have a shorter stage training time for a single stage than the state-of-the-art algorithm.

Face Deepfake-aware Recovery via Semantic-driven Facial Representation-based Watermarking

The Fortieth Annual Conference on Neural Information Processing Systems (NeurIPS)(Main Track), December 2026

Yuan-Chih Chen and Chun-Shien Lu

Yuan-Chih Chen Chun-Shien Lu

Abstract

Existing image watermarking methods typically entangle visual content with spatial structure, making them highly sensitive to geometric transformations and alignment discrepancies, especially in the presence of deepfake manipulations. In this paper, we propose a semantic-driven facial watermarking framework for robust identity recovery. The key idea is to decouple identity-related semantic information from spatial layout and encode it into a compact and spatially robust semantic representation. Specifically, we decompose a face into semantic components and aggregate deep features into component-wise latent representations, which are quantized via independent codebooks and converted into a compact bitstream for embedding. After decoding, the embedded semantic information is recovered and used to reconstruct identity-consistent facial content, even under slight geometric distortions and tampering. Experiments on CelebA-HQ and FFHQ demonstrate that our method significantly outperforms existing watermarking approaches in terms of reconstruction quality, identity preservation, and retrieval accuracy under both photometric and geometric attacks, validating the effectiveness of semantic component-wise encoding for reliable face recovery.

Position: Peer review should constrain evaluative authority

The Fortieth Annual Conference on Neural Information Processing Systems (NeurIPS) (Position Track), December 2026

Hanrui Wang, Timo Spinde, Chun-Shien Lu, and Isao Echizen

Hanrui Wang Chun-Shien Lu

Abstract

Our position shifts the target of peer-review reform from reducing uncertainty through evaluator improvement or judgment aggregation to constraining evaluative authority under persistent uncertainty. Scientific peer review aims to assure research quality but also allocates scarce academic opportunities, such as publication in high-impact venues. However, the review system is under growing pressure as submissions increase while reviewer capacity, expertise, scrutiny, and incentives cannot scale accordingly. As peer review operates under persistent evaluative uncertainty, evaluator improvement and judgment aggregation remain structurally limited. We therefore propose authority constrained peer review: before a consequential judgment is released as decision relevant, it should withstand structured scrutiny through critique-authority decoupling, independent secondary evaluation, contestability and reversibility, and accountability for unsupported influence. Our aim is to prevent weakly supported judgments from entering final deliberation with the same authority as judgments that have withstood structured scrutiny, while creating incentives for reviewers to justify consequential claims more carefully. For large-scale computer-science conferences, this position offers an institutional response to uneven review quality and reviewer-assignment dependence; more broadly, it provides a governance principle for any domain where uncertain expert judgments allocate scarce academic opportunities.

Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

Conference on Empirical Methods in Natural Language Processing (EMNLP 2026), October 2026

Kuan-Tang Huang, Cheng-Yeh Yang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, and Berlin Chen

Hsin-Min Wang

Abstract

Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translations and an acoustic anchor from speech; in principle, the mechanism can be applied to similar encoder--decoder multitask speech models. Through cross-modal adaptation, SAMA-ASR conditions decoder states on translation-derived semantic embeddings and a speech embedding, combining utterance-level meaning with speech-grounded evidence before token prediction. At evaluation time, these semantic anchors can be generated automatically by an upstream speech-to-text translator rather than supplied as oracle translations. Experiments on two 30-hour datasets covering the low-resource Sinitic varieties Taiwanese Hokkien and Hakka show that SAMA-ASR improves over acoustic, prior prompt-based, and semantic-only translation-guided baselines and remains effective in practical automatic semantic-anchor settings; translator-capacity analyses show that useful semantic anchors can be produced by a compact ST model.

CALM: Cycle-Aware Latent Modeling for Arteriovenous Fistula Assessment using Audio Foundation Models and Low-Level Descriptors

IEEE Journal of Translational Engineering in Health and Medicine, July 2026

Whenty Ariyanti, Ping-Yi Lin, Chia-Yu Hu, Shang-Feng Yang, Shin-Hung Tsai, Li-Ning Peng, Po-Hsun Huang, Hsin-Min Wang, and Yu Tsao

Hsin-Min Wang Yu Tsao

Abstract

Objective: Arteriovenous fistula (AVF) dysfunction is a major cause of morbidity in hemodialysis (HD) patients. Although conventional monitoring tools are clinically effective, they are often limited by cost, invasiveness, and operator dependence. This study proposes a cycle-aware, auscultation-based deep learning framework for objective and scalable assessment of AVF blood flow status. Methods and Procedures: We developed CALM, a physiologically informed learning framework that integrates Hilbert-based cycle segmentation with pretrained audio foundation model (AFM) representations. A clinical auscultation dataset (CHGH-AVF) from 188 HD patients at three standardized vascular access sites was collected and labeled using concurrent blood flow measurements. The proposed approach was evaluated on both three-class flow-severity classification (Low, Medium, and High) and binary flow-quality screening (Ideal vs. Unideal). Model variants incorporating domain-specific low-level descriptors and different feature fusion strategies were also investigated. Results: Across all auscultation sites, CALM consistently outperformed conventional handcrafted-feature baselines, deep learning baselines, and classical machine learning-based classifiers in patient-level flow-severity classification. Ablation studies demonstrated that physiologic cycle segmentation substantially improves discriminability, while integrating complementary low-level descriptors (LLDs) further enhanced performance across anatomically different recording locations. Although classification performance remained highest at the proximal sites, reduced performance at the venous site highlighted the influence of anatomical location and signal quality. Strong performance was also observed for binary flow-quality monitoring, particularly at the proximal auscultation sites. Conclusion: The proposed framework demonstrates the feasibility of AFM–based auscultation for non-invasive AVF flow assessment. By combining physiologic signal modeling with foundation model representations, CALM provides a robust and objective framework for automated assessment of vascular access function.Clinical Impact: This work introduces a multi-site clinical AVF auscultation dataset and establishes automated auscultation analysis as a low-cost and scalable adjunct to existing vascular access monitoring practices in HD care.Clinical and Translational Impact Statement: AFM-based auscultation has the potential to support objective screening of AVF dysfunction and may facilitate future integration into bedside or home monitoring systems, enabling more accessible and scalable vascular access monitoring.

Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations

IEEE Journal of Biomedical and Health Informatics, July 2026

Whenty Ariyanti, Kuan-Yu Chen, Sabato Marco Siniscalchi, Hsin-Min Wang, and Yu Tsao

Hsin-Min Wang Yu Tsao

Abstract

Perceptual voice quality assessment plays a vital role in diagnosing and monitoring voice disorders.Traditional methods, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and the Grade, Roughness, Breathiness, Asthenia, and Strain (GRBAS) scales, rely on expert raters and are prone to inter-rater variability, emphasizing the need for objective solutions. This study introduces the Voice Quality Assessment Network (VOQANet), a deep learning framework that employs an attention mechanism and Speech Foundation Model (SFM) embeddings to extract high-level features. To further enhance performance, we propose VOQANet+, which integrates self-supervised SFM embeddings with low-level acoustic descriptors—namely jitter, shimmer, and harmonics-to-noise ratio (HNR). Unlike previous approaches that focus solely on vowel-based phonation (PVQD-A), our models are evaluated on both vowel-level and sentence-level speech (PVQD-S) to assess generalizability. Experimental results demonstrate that sentence-based inputs yield higher accuracy, particularly at the patient level. Overall, VOQANet consistently outperforms baseline models in terms of root mean squared error (RMSE) and Pearson correlation coefficient across CAPE-V and GRBAS dimensions, with VOQANet+ achieving even greater performance gains. Additionally, VOQANet+ maintains consistent performance under noisy conditions, suggesting enhanced robustness for real-world and telehealth applications. This work highlights the value of combining SFM embeddings with low-level features for accurate and robust pathological voice assessment.

SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

ACM International Conference on Multimedia (ACM MM), November 2026

Kuan-Lin Chu, Jun-Cheng Chen, Chun-Shien Lu

Kuan-Lin Chu Chun-Shien Lu

Abstract

As speech generation models become increasingly realistic and widely accessible, concerns about the misuse, attribution, and governance of synthetic speech continue to grow.Watermarking provides a practical way to make synthesized speech traceable and verifiable. Most existing speech watermarking methods embed watermark information into signal-level representations, such as waveforms or spectrograms. Under sufficiently strong distortions, the embedded watermark may be weakened or destroyed, leading to degraded detectability. In this paper, we propose SSTMark, a training-free speech watermarking framework that operates at the semantic level through text watermarking. Unlike conventional signal-level watermarking methods, SSTMark encodes watermark information into the semantic content conveyed by generated speech, and detects the watermark from the recovered linguistic content. Experiments on AudioMarkBench demonstrate that SSTMark exhibits the strongest average robustness. Compared with the state-of-the-art baselines at a fixed false positive rate of 1%, SSTMark improves the average detection rate by 4.6% and 16.9% on signal-processing edits and compression edits, respectively.

Diffusion to Obfuscation: Time-Adaptive Synthesized Generation Against Gradient Leakage Attacks in Federated Learning

European Conference on Computer Vision (ECCV), September 2026

Farchan Hakim Raswa, Chun-Shien Lu, and Jia-Ching Wang

Chun-Shien Lu

Abstract

Recent studies demonstrate that Federated Learning (FL) is vulnerable to gradient leakage attacks (GLAs). Revisiting key insights in GLAs under FL reveals that (1) Defenses should focus on protecting semantic and fine-grained details of data; (2) GLAs are effective mainly in early rounds; and (3) To avoid semantic leakage, defenses shouldn’t infer the true labels of private images for obfuscation. Building upon these insights, we present a simple, time-adaptive defense strategy that obfuscates the private gradient by employing the gradient from a synthesized image. To this end, a client trains a diffusion model on its own private dataset to generate synthesized images that fit the distribution of private dataset but are distinct in fine-grained details. In addition, a synthesized image is generated conditioned on a non-identical label from the private image to resist semantic leakage. Defense analysis and empirical evaluations demonstrate that our time-adaptive and label agnostic method can better maintain the trade-off between privacy preservation and model utility against GLAs. Our implementation is publicly available at https://github.com/lalakitchen/Diff2Obs.

TiltDiff: Tilted Weight-Space Diffusion for Neural Network Generation

19th European Conference on Computer Vision (ECCV), September 2026

En-Ni Chuang, Hanjuan Huang, Hao-Jia Song, Hsing-Kuo Pao and Tyng-Luh Liu

Tyng-Luh Liu

Abstract

Generating functional neural-network weights from trained model collections is a central problem in weight-space learning. We introduce TiltDiff, a performance-tilted latent diffusion framework for neural-network generation. TiltDiff tokenizes network weights, encodes them into a compact latent space with a Transformer autoencoder, and uses a U-Net-based diffusion model to synthesize latent representations that decode into functional parameters. To favor stronger models, we weight the denoising loss by validation accuracy, biasing the learned distribution toward higher-performing regions of weight space. Experiments show that TiltDiff improves predictive performance, robustness to random parameter masking, and representational diversity over prior weight-generation methods. We further combine diffusion U-Net connectivity with attribution analysis to identify class-specific decision pathways. These pathways exhibit emergent correspondence across independently generated models despite differing raw parameters, and pathway-level masking verifies their importance for target-class prediction. Our results show that performance-tilted diffusion generates accurate, robust, diverse, and structurally interpretable neural-network weights.

Creat3r: Confidence Reaggregation for Exploration-aware Active 3D Reconstruction

The 43rd International Conference on Machine Learning (ICML), July 2026

Chih-Jung Tsai, Hwann-Tzong Chen and Tyng-Luh Liu

Tyng-Luh Liu

Abstract

We present Creat3r, an iterative next-best-view (NBV) selection framework for efficient, high-quality 3D reconstruction. Starting from a small seed set of image--pose pairs, Creat3r repeatedly selects the most informative next camera pose. After each pose is chosen, the corresponding image is acquired and added to the multi-view set to update a 3DGS reconstruction. To guide selection, Creat3r constructs an intermediate point cloud and estimates reconstruction reliability via a novel \emph{3D confidence field}, which is projected to candidate poses through Gaussian projection to produce \emph{2D confidence} and \emph{exploration maps}. These maps balance exploitation of reliable regions and exploration of uncertain or unseen areas under computational constraints. Experiments with standard 3DGS show that Creat3r consistently outperforms baselines in novel view synthesis and surface reconstruction, achieving higher SSIM and F1 scores with fewer views.

DNA-DETR: sequence representation matters in object detection for functional genomic elements

Briefings in Bioinformatics, June 2026

Tsai, B.S., Wong, J.Y., and Tsai, H.K.*

Tsai B.S. Wong

Abstract

Object detection has revolutionized multiple domains by enabling models to jointly classify and localize targets within data. Yet, its potential in genomic sequence analysis remains largely unexplored. Here, we introduce DNA-DETR, an adaptation of the DETR architecture for one-dimensional genomic object detection. Surprisingly, the direct application of object detection to DNA sequences yielded poor performance, even for elements with simple definitions such as Non-B DNA. We found that the widely used one-hot encoding failed to capture key structural features of several Non-B DNA types. To address this limitation, we systematically investigated how different sequence representations, including one-hot encoding, dot matrix, and their combination, affect detection accuracy and model generalization. Our experiments demonstrate that the choice of representation profoundly affects both localization and classification. Notably, the combined representation consistently outperformed single representations, particularly for complex sequence elements. Our findings suggest that there is no universal “one-representation-fits-all” solution in sequence feature learning. Despite the common perception that end-to-end learning diminishes the importance of representation, our results highlight that thoughtful selection of sequence representation remains critical for model design.

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

IEEE Signal Processing Letters, June 2026

Chien-Chun Wang,Hung-Shin Lee,Hsin-Min Wang, and Berlin Chen

Hsin-Min Wang

Abstract

Evaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS). Most automatic MOS estimators are trained with point-wise regression or distributional classification. These objectives do not directly optimize rank-based metrics and provide weak geometric constraints for cross-modal coherence. To address these gaps, we propose DeRA-MOS, a decoupled optimization framework for TTM evaluation. For MI, we introduce a batch-aware listwise ranking loss that models relative order within each mini-batch and better aligns with evaluation based on Spearman's rank correlation coefficient (SRCC). For TA, we introduce a score-anchored modality alignment loss that maps human scores to target audio-text similarity and regularizes the latent space before fusion. By effectively mitigating the point-wise training mismatch and modality drift, experiments on MusicEval demonstrate that our decoupled framework yields substantial improvements in both MI and TA ranking metrics, establishing a robust paradigm for large-scale TTM evaluation.

DIffUMI: Training-Free Universal Model Inversion via Unconditional Diffusion for Face Recognition

IEEE Transactions on Information Forensics and Security , April 2026

Hanrui Wang, Shuo Wang, Chun-Shien Lu, and Isao Echizen

Chun-Shien Lu

Abstract

Face recognition poses serious privacy risks due to its reliance on sensitive and immutable biometric data. While modern systems mitigate privacy risks by mapping facial images to embeddings (commonly regarded as privacy-preserving), model inversion attacks reveal that identity information can still be recovered, exposing critical vulnerabilities. However, existing attacks are often computationally expensive and lack generalization, especially those requiring target-specific training. Even training-free approaches su er from limited identity controllability, hindering faithful reconstruction of nuanced or unseen identities. In this work, we propose Di MI, the first di usion-driven, training-free model inversion attack. Di MI introduces a novel pipeline combining robust latent code initialization, a ranked adversarial refinement strategy, and a statistically grounded, confidence-aware optimization objective. Di MI applies directly to unseen target identities and face recognition models, o ering greater adaptability than trainingdependent approaches while significantly reducing computational overhead. Our method achieves 84.42%–92.87% attack success rates against inversion-resilient systems and outperforms the best prior training-free GAN-based approach by 4.01%–9.82%. The implementation is available at https://github.com/azrealwang/ Di MI.

Harnessing Sequence Embedding and Ensemble Learning to Identify Antifungal Peptides with Low Hemolytic Risk

ACS OMEGA, April 2026

Chung-Yen Lin,Wen-Chih Cheng, U-Lin Chen, Tzu-Tang Lin, Li-Hang Hsu, Yang-Hsin Shih, I-Hsuan Lu, Ying-Lien Chen, Shu-Hwa Chen

Chung-Yen Lin Shu-Hwa Chen Wen-Chih Cheng I-Hsuan Lu

Abstract

The increasing prevalence of fungal infections represents a growing threat to human health, driven in part by the misuse of antibiotics and the rising incidence of resistance to conventional antifungal agents. Antifungal peptides (AFPs) have emerged as promising alternatives due to their diverse mechanisms of action and their relatively low propensity to develop resistance. To facilitate the systematic discovery of AFPs, we developed AI4AFP. This computational framework integrates curated antifungal peptide resources with advanced machine learning approaches to predict antifungal potential directly from peptide sequences.

Using a comprehensive dataset, we constructed a seven-model ensemble that combines multiple sequence encoding strategies, including ProtBERT-BFD, PC6, and Doc2Vec, with diverse learning algorithms, including random forests, support vector machines, convolutional neural networks, and fine-tuned BERT models. This ensemble demonstrated robust performance on an independent test set, achieving 0.94 in accuracy and 0.89 in Matthews correlation coefficient, outperforming existing AFP prediction methods. Importantly, the predicted AFP score is intended to reflect the general antifungal potential rather than species-specific potency.

Experimental validation against representative fungal pathogens, including Candida albicans, Candida glabrata, and Cryptococcus neoformans, revealed that peptides with high predicted AFP scores exhibited context-dependent antifungal activity. Several candidates displayed pronounced inhibitory effects against specific species, despite limited activity against others, highlighting the inherent species-dependence of antifungal efficacy and supporting the role of AI4AFP as a prioritization tool rather than a species-specific predictor.

To complement antifungal prediction, we further developed a hemolysis classifier that incorporates both peptide sequence and applied concentration as continuous inputs, enabling explicit modeling of the dose-dependent nature of hemolytic toxicity. Experimental determination of the minimum concentration inducing 10% hemolysis (MHC₁₀) provided an empirical safety reference, allowing antifungal activity to be interpreted alongside concentration-dependent toxicity. All models and validation results are implemented on a user-friendly web server, AI4AFP (https://axp.iis.sinica.edu.tw/AI4AFP), providing an accessible platform for the discovery and prioritization of antifungal peptides, with consideration of both efficacy and safety.

Regret-Guided Search Control for Efficient Learning in AlphaZero

the Fourteenth International Conference on Learning Representations (ICLR), April 2026

Yun-Jui Tsai, Wei-Yu Chen, Yan-Ru Ju, Yu-Hung Chang, Ti-Rong Wu

Yan-Ru Ju Yu-Hung Chang Ti-Rong Wu

Abstract

Reinforcement learning (RL) agents achieve remarkable performance but remain far less learning-efficient than humans. While RL agents require extensive self-play games to extract useful signals, humans often need only a few games, improving rapidly by repeatedly revisiting states where mistakes occurred. This idea, known as search control, aims to restart from valuable states rather than always from the initial state. In AlphaZero, prior work Go-Exploit applies this idea by sampling past states from self-play or search trees, but it treats all states equally, regardless of their learning potential. We propose Regret-Guided Search Control (RGSC), which extends AlphaZero with a regret network that learns to identify high-regret states, where the agent's evaluation diverges most from the actual outcome. These states are collected from both self-play trajectories and MCTS nodes, stored in a prioritized regret buffer, and reused as new starting positions. Across 9x9 Go, 10x10 Othello, and 11x11 Hex, RGSC outperforms AlphaZero and Go-Exploit by an average of 77 and 89 Elo, respectively. When training on a well-trained 9x9 Go model, RGSC further improves the win rate against KataGo from 69.3% to 78.2%, while both baselines show no improvement. These results demonstrate that RGSC provides an effective mechanism for search control, improving both efficiency and robustness of AlphaZero training. Our code is available at https://rlg.iis.sinica.edu.tw/papers/rgsc.