-
NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild
http://arxiv.org/abs/2604.11487v1 · arxiv
This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE workshop at CVPR 2026. The goal of this challenge was to develop detection models capable of distinguishing real images from generated ones in realistic
-
DeepTest Tool Competition 2026: Benchmarking an LLM-Based Automotive Assistant
http://arxiv.org/abs/2604.12615v1 · arxiv
This report summarizes the results of the first edition of the Large Language Model (LLM) Testing competition, held as part of the DeepTest workshop at ICSE 2026. Four tools competed in benchmarking an LLM-based car manual information retrieval application, with the objective of identifying user inp
-
Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs
http://arxiv.org/abs/2605.10877v1 · arxiv
Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, whic
-
RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
http://arxiv.org/abs/2605.09568v3 · arxiv
RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an E
-
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
http://arxiv.org/abs/2603.08092v1 · arxiv
Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT
-
DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task
http://arxiv.org/abs/2610.06298v1 · arxiv
Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and
-
The Eleventh NTIRE 2026 Efficient Super-Resolution Challenge Report
http://arxiv.org/abs/2604.03198v1 · arxiv
This paper reviews the NTIRE 2026 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a network that reduces one or several aspects, such as runtime, parameters, and FLOPs, while maintaining PSNR of around 26
-
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
http://arxiv.org/abs/2605.02695v1 · arxiv
SemEval-2026 Task 9 is focused on multilingual polarization detection. Specifically, it covers the identification of multilingual, multicultural and multievent polarization along three axes (in subtasks), namely detection, type, and manifestation. Online polarization presents a concern, because it i
-
NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning
http://arxiv.org/abs/2607.16603v1 · arxiv
This paper presents the methodologies and results of the NOWJ team's participation across all five tasks of the COLIEE 2026 competition. For Task 1 (Legal Case Retrieval), we propose a four-stage pipeline comprising candidate filtering, dense retrieval with complementary embedding models, cross-enco
-
mdok-style at SemEval-2026 Task 10: Finetuning LLMs for Conspiracy Detection
http://arxiv.org/abs/2605.02712v1 · arxiv
SemEval-2026 Task 10 is focused on conspiracy detection. Specifically, the goal is to detect whether a Reddit comment expresses a conspiracy belief. Our submitted mdok-style system utilizes data augmentation and self-training (to cope with a rather small amount of training data) to finetune the Qwen
10 hasil · 797 md