Waktu: Semua Jam Hari Minggu Bulan Tahun  ·  Bahasa: Indonesia Semua
  1. Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

    http://arxiv.org/abs/2607.09623v1 · arxiv

    We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under

  2. BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes

    http://arxiv.org/abs/2607.24137v1 · arxiv

    This paper describes the BioSentinel team's participation in EXIST 2026 Task 2.2: Source Intention in Memes, part of the CLEF 2026 evaluation campaign. The task requires classifying the communicative intent behind memes as direct, judgemental, or no (non-sexist), under a Learning with Disagreement (

  3. AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task

    http://arxiv.org/abs/2606.03967v1 · arxiv

    We describe AlignAtt4LLM, an IWSLT 2026 simultaneous speech translation system for English to German, Italian, and Chinese. The system is a synchronous cascade: Qwen3-ASR with forced alignment produces an incrementally updated source transcript, and Gemma-4 E4B-it translates that prefix under an MT-

  4. NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild

    http://arxiv.org/abs/2604.11487v1 · arxiv

    This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE workshop at CVPR 2026. The goal of this challenge was to develop detection models capable of distinguishing real images from generated ones in realistic

  5. DeepTest Tool Competition 2026: Benchmarking an LLM-Based Automotive Assistant

    http://arxiv.org/abs/2604.12615v1 · arxiv

    This report summarizes the results of the first edition of the Large Language Model (LLM) Testing competition, held as part of the DeepTest workshop at ICSE 2026. Four tools competed in benchmarking an LLM-based car manual information retrieval application, with the objective of identifying user inp

  6. Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs

    http://arxiv.org/abs/2605.10877v1 · arxiv

    Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, whic

  7. RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations

    http://arxiv.org/abs/2605.09568v3 · arxiv

    RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an E

  8. DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task

    http://arxiv.org/abs/2610.06298v1 · arxiv

    Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and

  9. Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge

    http://arxiv.org/abs/2603.08092v1 · arxiv

    Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT

  10. mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection

    http://arxiv.org/abs/2605.02695v1 · arxiv

    SemEval-2026 Task 9 is focused on multilingual polarization detection. Specifically, it covers the identification of multilingual, multicultural and multievent polarization along three axes (in subtasks), namely detection, type, and manifestation. Online polarization presents a concern, because it i

10 hasil · 1109 md