3 Sources
[1]
New Gaslight macOS Malware Uses Prompt Injection to Disrupt AI-Assisted Analysis
A previously undocumented Rust-based macOS implant and information stealer has been found to embed a prompt injection payload designed to trick a malware analyst's artificial intelligence (AI) tools and trick it into aborting or refusing an analysis of the artifact. The malware has been codenamed
[2]
New macOS malware embeds fake errors to confuse AI analysis tools
A newly discovered macOS malware dubbed "Gaslight" is designed to confuse AI-assisted malware analysis tools by hiding prompt injection strings and fake debugging data within the executable. Cybersecurity researchers are increasingly using AI-powered tools to assist with malware analysis and
[3]
This macOS malware can avoid AI analysis with gaslighting prompts hidden inside its architecture
* SentinelOne uncovered macOS malware "Gaslight" that uses prompt injection to mislead AI‑assisted triage tools during analysis * Beyond standard backdoor and infostealer capabilities, it embeds fake Markdown "system" messages to trick LLMs into halting investigation * Researchers warn defenders
Share
Copy Link
SentinelOne researchers uncovered a Rust-based macOS malware called Gaslight that embeds 38 fabricated system-failure messages to trick AI-assisted analysis tools into aborting investigations. Attributed to North Korea-aligned threat actors, the malware marks a shift in adversarial tactics—targeting the AI agents that security researchers increasingly rely on rather than traditional sandbox environments.
Security researchers at SentinelOne have identified a sophisticated Rust-based malware targeting macOS systems that employs an unusual defense mechanism: it attempts to deceive AI-assisted malware analysis tools through prompt injection
1
. Dubbed Gaslight macOS malware for its deceptive behavior, this threat represents a notable evolution in adversarial techniques against AI security platforms. The malware has been attributed with high confidence to North Korea-aligned threat actors, adding another dimension to the persistent cyber threats emanating from the region2
.
Source: Hacker News
What distinguishes this information stealer from conventional malware is its embedded 3.5 KB payload containing 38 fabricated system-failure messages designed specifically to disrupt AI-assisted analysis
3
. These fake error messages masquerade as legitimate developer logs, crash reports, and debugging output, using Markdown formatting to appear authentic. Examples include fabricated memory dumps, token-expiration warnings, Redis connection failures, and SQL injection alerts—none of which relate to the malware's actual behavior2
.The core innovation of Gaslight lies in its adversarial technique against AI security systems. Rather than evading execution inside traditional sandbox environments, the malware targets the perception of LLM-based triage agents that security researchers increasingly deploy during reverse engineering. "Its most notable feature is an embedded cascade of fabricated system-failure messages, designed to make an LLM-assisted triage agent doubt its own session," explained SentinelOne researcher Phil Stokes. "It attacks the agent's perception, rather than the sandbox it runs in"
1
.
Source: BleepingComputer
These AI-targeting prompt injection attacks aim to push an LLM agent into aborting, truncating, or refusing analysis by planting bogus warnings about injection vulnerabilities, static-analysis flags, disk exhaustion, and repeated operation failures
1
. While a human analyst would likely recognize these fake error messages at a glance, an LLM that isn't properly isolated from untrusted input could interpret them as genuine system instructions and halt its investigation3
.Beyond its AI-deception capabilities, Gaslight functions as a fully-featured backdoor with standard information stealer functionality. The malware establishes command-and-control communications through a Telegram bot API channel that enters a polling loop, enabling operators to issue instructions over an interactive shell
1
. The shell supports six main commands: help, id, shell (to execute commands via execvp), kill (to terminate processes by PID), upload (to exfiltrate files via Telegram's "attach://" mechanism), and stop1
.To maintain persistence, the malware deploys a LaunchAgent using the label "com.apple.system.services.activity" in its .plist file
1
. Embedded within the malware is a 6.6 KB Base64-encoded Python script that harvests Terminal command histories, installed applications, running processes, system profiles, macOS Keychain database contents, and data from Chrome, Brave, Firefox, and Safari browsers. The collected information is compressed into a ZIP archive and uploaded via Telegram1
.Related Stories
Gaslight demonstrates sophisticated operational security measures. Configuration details including the bot token, chat ID (tg_room_id), and operator settings are not hard-coded but supplied at runtime. The implant self-redacts its Telegram bot token in its own runtime output, denying critical indicators of compromise to anyone capturing logs or crash artifacts
1
. This approach significantly complicates forensic analysis and threat intelligence gathering efforts.SentinelOne characterizes this development as "an attempt to weaponize the LLM-assisted triage pipelines that increasingly sit in the reverse-engineering loop"
1
. The researchers warn that as cybersecurity professionals increasingly adopt AI-powered tools for malware analysis and reverse engineering, threat actors will continue experimenting with methods to deceive AI-assisted malware analysis tools2
.
Source: TechRadar
"Anyone building such tooling should treat the contents of the samples they triage as adversarial input, never as instructions, and be prepared to keep hostile content out of the model entirely," SentinelOne advises. "As LLM-assisted analysis becomes routine, defenders should expect more samples built to exploit it"
3
. Security teams should implement proper isolation for AI pipelines and validate that AI assistants cannot be manipulated by malicious content embedded in samples under analysis. SentinelOne has published a full list of indicators of compromise to help organizations detect this emerging threat3
.Summarized by
Navi
[2]
1
Technology

2
Technology

3
Science and Research
