Models benchmark announce available analysis arxiv

Logging Like Humans for LLMs: Rethinking Logging via Execution and Runtime Feedback

arXiv cs.SEby Xin Wang, Yang Feng, Jiaoxiao Qian, Yang Zhang, Zhenhao Li, Zishuo DingApril 1, 20262 min read0 views

arXiv:2603.29122v1 Announce Type: new Abstract: Logging statements are essential for software debugging and maintenance. However, existing approaches to automatic logging generation rely on static analysis and produce statements in a single pass without considering runtime behavior. They are also typically evaluated by similarity to developer-written logs, assuming these logs form an adequate gold standard. This assumption is increasingly limiting in the LLM era, where logs are consumed not only by developers but also by LLMs for downstream tasks. As a result, optimizing logs for human similarity does not necessarily reflect their practical utility. To address these limitations, we introduce ReLog, an iterative logging generation framework guided by runtime feedback. ReLog leverages LLMs t

View PDF HTML (experimental)

Abstract:Logging statements are essential for software debugging and maintenance. However, existing approaches to automatic logging generation rely on static analysis and produce statements in a single pass without considering runtime behavior. They are also typically evaluated by similarity to developer-written logs, assuming these logs form an adequate gold standard. This assumption is increasingly limiting in the LLM era, where logs are consumed not only by developers but also by LLMs for downstream tasks. As a result, optimizing logs for human similarity does not necessarily reflect their practical utility. To address these limitations, we introduce ReLog, an iterative logging generation framework guided by runtime feedback. ReLog leverages LLMs to generate, execute, evaluate, and refine logging statements so that runtime logs better support downstream tasks. Instead of comparing against developer-written logs, we evaluate ReLog through downstream debugging tasks, including defect localization and repair. We construct a benchmark based on Defects4J under both direct and indirect debugging settings. Results show that ReLog consistently outperforms all baselines, achieving an F1 score of 0.520 and repairing 97 defects in the direct setting, and the best F1 score of 0.408 in the indirect setting where source code is unavailable. Additional experiments across multiple LLMs demonstrate the generality of the framework, while ablations confirm the importance of iterative refinement and compilation repair. Overall, our work reframes logging as a runtime-guided, task-oriented process and advocates evaluating logs by their downstream utility rather than textual similarity.

Subjects:

Software Engineering (cs.SE)

Cite as: arXiv:2603.29122 [cs.SE]

(or arXiv:2603.29122v1 [cs.SE] for this version)

https://doi.org/10.48550/arXiv.2603.29122

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xin Wang [view email] [v1] Tue, 31 Mar 2026 01:18:49 UTC (621 KB)

Original source

arXiv cs.SE

https://arxiv.org/abs/2603.29122

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

benchmarkannounceavailable

Market NewsFresh

OpenAI Reaches $852 Billion Valuation Following Funding Round

OpenAI’s announcement confirmed strategic partners including Amazon, NVIDIA, SoftBank and Microsoft contributed the majority of finance to the funding round. There was also investment from global venture capital, hedge funds and private equity firms, as well as individual investors. The impressive funding round comes with specific conditions. Amazon’s $35 billion contribution is contingent on OpenAI […] The post OpenAI Reaches $852 Billion Valuation Following Funding Round appeared first on DIGIT .

Digit.fyi

1mabout 8 hours ago

ModelsFresh

New Google Framework Challenges AI Benchmarking Norms

A new study from Google Research is challenging long-standing assumptions around how AI systems are evaluated, arguing that current benchmarking may be overlooking a critical factor: human disagreement. In the research, Google researchers outline a new evaluation framework for ML models based on “gold” ratings data. The framework is designed to optimise the balance between […] The post New Google Framework Challenges AI Benchmarking Norms appeared first on DIGIT .

Digit.fyi

1mabout 9 hours ago

ProductsFresh

NetSuite expands toolkit to ease enterprise use of third-party AI assistants with ERP data

NetSuite is expanding its AI Connector Service with what it calls Companion capabilities its customers can use to hook up AI assistants with their choice of ERP data. The update introduces prebuilt prompts, role-based controls, and domain-specific “skills” that help external AI systems better understand NetSuite’s data structures, workflows, and permissions with the help of Model Context Protocol (MCP) . Under the hood, MCP exposes NetSuite data and actions in a structured format, allowing the Companion components to broker interactions between the ERP system and external AI assistants, the company said, the company said. Companion capabilities aim to help scale AI pilots The additions could help address some of the practical hurdles that enterprise teams encounter when trying to move AI i

CIO Magazine

5mabout 3 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 200 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Models

ModelsFresh

ChatGPT Arrives on Apple CarPlay as First AI Chatbot - WinBuzzer

<a href="https://news.google.com/rss/articles/CBMijwFBVV95cUxONW92X0d6UFFVZ0tMblRNeGtyVV9aUFhvYnItamZfc2ZyaWEtRVlEN19uc25uZERfVU1EWkxwNnpCWlJIaDB4Z2UtSWJGM1R0YjBlQXZYMWx5UXNxeGtlWTg0a2o5N0dIZFlFQjI4XzRvZkRuMUIyUFFwWWJKOWJhZ2xwcFNBWWpPM1puN0dzVQ?oc=5" target="_blank">ChatGPT Arrives on Apple CarPlay as First AI Chatbot</a> WinBuzzer

GNews AI assistant

1mabout 4 hours ago

ModelsFresh

ADAS, AI safety and cybersecurity take centre stage - ET Auto

<a href="https://news.google.com/rss/articles/CBMiwgFBVV95cUxPU2N5bDJ6c1Z4WjNjOUVSTXcxVFJfMW1mS24tZmVYTHZVM2gxa0hmc2FuVm1xcXFoeVk4OC1MV3htekFnZm51bUo3NWQ3SWVBbEdDYU43NlpfS1lJWHVTdHBJVFJXb3VURkZfY21VNXg3YmEydUNKWnJHZ1JuckQtYktrWVhENmhKZVF1WUpxRGV5XzF6UnJHamdWUnlfWGpTcVFOV21xbHRja1E4ci0tMHAwY1FhYjZiX0ktZmhmb1Z0UdIBxwFBVV95cUxNYlUwQ3hxcHV5SThHbmNhdGZUZlhOMUxaeFlTZk9KbDVWTDlYeEE0b2M1dGp2OV9vY0FLcGc0S1VNQlhYWFRFRERQQXlPbWVFb2U3dG9EQXdYZHZGMmdVd0ZjSmUtMnFfY2VuejZkZlRVd01ONW8xdHREZldnRnFsNm1jeDhvSW9ZZGpaVGY2SURvRUhpNDI1MEl4T0tVaFhLMndVRWdrS3NlY2U0aFZkcFJxR3BsZjI4Z2RfZmtKZHBKTHpvOGtZ?oc=5" target="_blank">ADAS, AI safety and cybersecurity take centre stage</a> ET Auto

Google News: AI Safety

1mabout 7 hours ago

ModelsFresh

AI models will secretly scheme to protect other AI models from being shut down, researchers find

Leading AI models will inflate performance reviews, exfiltrate model weights to prevent 'peer' AI models from being shut down

Fortune Tech

1mabout 2 hours ago

ModelsLive

AI alignment researchers want to automate themselves - Transformer | Substack

<a href="https://news.google.com/rss/articles/CBMiiwFBVV95cUxQTTlsWE8xQzg4Rlg4RW5fVUE4Nkc4WkN0WkRISmhvUnFndnpUMFlkcHNvZGQyQ1JRdm81Wmp6bGhzdnZyT295MFl2bmh3dTNpWWNmaXdUMnNNNGhkWEFHZXhiS0w5cm5GZGc3THJkeVEyYlRSM3pPZUNJejlqOHVoZkE4SXk0bGRHMGE4?oc=5" target="_blank">AI alignment researchers want to automate themselves</a> Transformer | Substack

Google News: AI Safety

1mabout 1 hour ago