← heapsort-ai

uncertainty

6 items

RESEARCHarXiv CS.AI·19d ago

$ECUAS_n$: A family of metrics for principled evaluation of uncertainty-augmented systems

This research proposes a new family of metrics, $ECUAS_n$, for evaluating uncertainty-augmented (UA) systems in automated decision-making. It argues that existing evaluation approaches are insufficient for assessing overall performance of UA systems, where predictive uncertainty is crucial for users to make informed decisions.

30
RESEARCHarXiv CS.CL·25d ago

When Evidence Conflicts: Uncertainty and Order Effects in Retrieval-Augmented Biomedical Question Answering

This research evaluates large language models (LLMs) in biomedical question answering, specifically addressing their reliability when faced with conflicting or incomplete evidence. It reveals that LLM accuracy significantly drops, and predictions flip, when the order of correct and contradictory documents is reversed, highlighting issues with order effects and the need for conflict-aware abstention.

27
RESEARCHarXiv CS.AI·8d ago

Uncertainty-Aware and Temporally Regulated Expert Advice in Reinforcement Learning for Autonomous Driving

This paper proposes an uncertainty-aware framework for reinforcement learning in autonomous driving, leveraging expert advice to guide exploration safely while avoiding long-term dependence. It employs adaptive thresholds for advice triggering and a commitment-cooldown strategy to regulate guidance, demonstrating improved performance in CARLA simulations.

27