Workshop on Insights from Negative Results in NLP
Budapest, Hungary, October 29, 2026
(co-located with EMNLP)
Geometric Separation Without Functional Specialization: Auditing Subspace Disentanglement in Dual-Encoder LLM Paraphrase Detection
Govind Arun, Ashish Abraham and Gadha Lekshmi P
Exploring Dowker Homology for Sentence Similarity
Marius Huber and Juri Opitz
Does Pragmatic Transfer Help Euphemism Detection? Evidence from Target-Span Marking
Whitney Poh, Noriko Takahashi, Patrick Lee, Julia Sammartino, Libby Barak, Jing Peng and Anna Feldman
Risk-Averse Failure Modes of Jensen-Shannon Distance in RLVR Distributional Alignment
Sebastian Loftus, Xinpeng Wang, Marie-Catherine De Marneffe and Barbara Plank
Controlled Fine-Tuning for Two-Tower Recommenders: Why Gradual Unfreezing Can Fail and Cosine Regularization Suffices
Gabriel Dershowitz and Djallel Bouneffouf
Do Soft Prompts Have Useful Wavelet Structure? Negative Evidence In A Controlled Setup
Soumyadeep Dey, Srikanth Tenneti, Bamdev Mishra, Shobhit Niranjan and Mayur Datar
Decomposing LLM-Judge–Human Agreement: How Much Is Inherited Surface Preference?
Gaurav Kumar
The Flip-Test: A Behavioral Diagnostic for Multi-Agent Debate Failures
Julia Hu, Alfred Shen and Kumar Lakshmipathi
Activation-Based Active Learning for In-Context Learning: Challenges and Insights
Yaseen Osman, Geoff Merrett and Stuart Middleton
The Architecture Was Doing the Work: Why NLI Debiasing Results Do Not Generalise
Prabhjot Singh
When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA
Yingrui Li and Han Chen
No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study
Gürkan Soykan and Gözde Şahin
Short-Context Gains, Long-Context Drift: Negative Results for Token-Conditioned RoPE
Allan Kazakov
When Gradient Importance Lies: Adaptive LoRA Rank Allocation Fails Under GRPO
Yash Sawant
Do Not Score the Edited Answer: A Negative Result on Loose Evaluation for Verifiable Instruction Following
José Luis Delgado
Perturbation, Not Provenance: Why Activation Steering Does Not Make a Robust Text Watermark
Sergey Pletenev