Files
hapi/web/src/lib/remark-non-https-autolink.ts
T
Junmo KimandGitHub 6e32b20524 feat(web): linkify custom URI schemes with a confirm prompt (#633)
* feat(web): add UriConfirmDialog component

Add a Radix Dialog-based confirmation modal for custom URI scheme
navigation. Follows the RenameSessionDialog pattern.

- UriConfirmDialog: shows URI, scheme label, Cancel/Open/Always-allow buttons
- i18n keys: dialog.uri.{title,description,open,alwaysAllow}

* feat(web): autolink non-https URI schemes in markdown

Add a remark plugin that converts raw `scheme://...` text nodes into
link nodes for non-http(s) schemes. GFM already handles http/https;
this plugin handles the remainder (obsidian://, vscode://, slack://, etc.).

- No scheme allowlist: every `scheme://` pattern is converted; the
  sanitize layer (urlTransform) and onClick layer (classifyScheme) handle
  blocking/confirmation downstream.
- Runs before remarkStripCjkAutolink so the CJK-strip plugin sees the
  new link nodes and can trim trailing CJK punctuation from them.
- Trailing punctuation (.,;!?) stripped from matched URIs.
- Unit tests: conversion, partial-match, escape, explicit link bypass,
  code-block bypass, trailing-punct trimming.

* feat(web): linkify custom URI schemes via markdown <a> handler

Wire up 4-layer URI security policy in the markdown renderer:

1. URL sanitize (deny-only): urlTransform strips javascript:/data:/vbscript:/file:
   using classifyScheme as single source of truth (handles percent-encoding,
   case-insensitive, whitespace-prefix bypass patterns).

2. onClick intercept: custom <A> component classifies each href —
   - IANA safe (https/http/irc/ircs/mailto/xmpp): navigate directly.
   - Deny (javascript/data/vbscript/file): preventDefault silently.
   - Custom (obsidian/vscode/slack/…): preventDefault + open UriConfirmDialog.

3. UriConfirmProvider: one dialog lifted to each markdown root (MarkdownText,
   Reasoning, MarkdownRenderer). Shared isAllowed state across all <a> tags in
   the subtree — "Always allow" click updates every link in one React commit.

4. Intra-tab cross-provider sync (P7e.1): module-level schemeListeners Set so
   sibling UriConfirmProviders (MarkdownText + Reasoning in AssistantMessage)
   receive allowed-scheme updates synchronously without waiting for the window
   storage event (which only fires in other tabs). Cross-tab sync continues via
   the existing window storage event listener.

5. "Always allow" persisted to localStorage (hapi-allowed-schemes). Custom
   schemes once allowed navigate directly on subsequent clicks, no dialog gate.
   href="#" in DOM for unallowed custom schemes prevents middle-click bypass.
   Deny-scheme href="" prevents any navigation even if localStorage tampered.

Security: classifyScheme decodes percent-encoding before scheme extraction,
blocking %6Aavascript:, jav%61script:, javascript%3A (single-encoded colon)
and double-encoded variants. DENY_SCHEMES checked after localStorage lookup so
tampered allowed-list cannot promote deny schemes.

Tests: classifyScheme 6-axis security bypass, denyOnlyTransform, localStorage
roundtrip, cross-tab storage event, <A> click handler cases.

* fix(web): block control-char-spliced deny schemes in classifyScheme

Browsers silently strip ASCII control characters (\t, \n, \r) and
whitespace from URL scheme names during navigation. A scheme like
`java\nscript:alert(1)` was navigated as `javascript:` while our
literal string comparison classified it as 'custom', allowing it
past the deny list and into window.open().

Introduce normalizedScheme() that:
- applies 2 rounds of decodeURIComponent so double-encoded schemes
  (javascript%253A → javascript%3A → javascript:) are fully unwrapped
  before comparison
- strips [\x00-\x1F\x7F\s] from the extracted scheme name, matching
  the browser's own normalization

classifyScheme() now delegates to normalizedScheme() so both the
denyOnlyTransform (urlTransform) path and the <A> onClick path benefit
from the same normalization.

Tests added for \n / \t / \r / space spliced into scheme, and verify
that double-encoded colon is now caught via scheme-match (not just
the no-colon fallback).

* fix(web): preserve relative markdown links from being blocked

Relative / no-scheme hrefs (/settings, ./foo, #section, ?q=1) were
silently preventDefault'd in <A>'s onClick handler. denyOnlyTransform
correctly passed them through (no colon → not a scheme URL), but the
click handler called classifyScheme(href) which returned 'deny' for
any input with no valid scheme separator — then the deny branch fired.

Add hasScheme(href): checks whether the first ':' appears before any
path/query/fragment boundary ('/', '?', '#'). When hasScheme is false
the href is treated as 'iana' so the browser or SPA router can navigate
normally with no dialog and no preventDefault.

Also wrap renderA() with <I18nProvider> so the UriConfirmDialog that
UriConfirmProvider may render does not throw outside its translation
context during tests.

Fixes a regression that broke all relative-path markdown links once the
custom-URI-scheme onClick handler was added.

* test(web): cover percent-encoded scheme control char + protocol-relative href

Round-5 internal hostile review noted two coverage gaps on the bot-fixup commits:

- `java%0Ascript:alert(1)` (percent-encoded newline in the scheme name) takes the
  same decode→strip code path as the literal `java\nscript:` case but was only
  tested literally. Add an explicit test so a future refactor that drops the
  decode-then-strip ordering would be caught.
- Protocol-relative URLs (`//host/path`) have no colon, so `hasScheme` returns
  false and `<A>` treats them as scheme-less — browsers then navigate them as
  the current origin's protocol. Existing relative-href tests covered absolute
  paths, hashes, queries, and colon-in-path, but not the protocol-relative
  variant. Add one assertion.

Also extend the `hasScheme` JSDoc to note that protocol-relative URLs are
intentionally treated as scheme-less.

* fix(web): preserve balanced parens/brackets in autolinked URIs

The trailing-punctuation strip used to drop every `)` / `]` from the end
of a matched URI, even when the URL body had an unmatched opener. So a
URI like `obsidian://open?file=Note(1)` was rendered with href
`obsidian://open?file=Note(1` plus a separate `)` text node, opening a
broken deep link.

Match the GFM autolink-literal behaviour: when the trailing character is
`)` or `]`, keep it iff the URL body has more opening counterparts than
closers (so the trailing closer balances an earlier opener and belongs
to the URL). Other trailing punctuation (`.,;!?:>'"`) and unmatched
closers still strip as before.

Add tests for the balanced cases (`Note(1)`, `Note[1]`, nested
`(a(b)c)`), the "balanced URL followed by a period" case, and a
regression test that an unmatched `).` after a URL is still stripped.
2026-05-18 10:40:32 +08:00

174 lines
5.8 KiB
TypeScript

/**
* Remark plugin that converts raw non-https URI scheme text into link nodes.
*
* GFM (`remark-gfm`) already handles `http://`, `https://`, and `www.` autolinks.
* This plugin handles the remainder: any `scheme://...` pattern where the scheme
* is NOT `http` or `https` (to avoid duplicating GFM's work).
*
* Pipeline position: before `remarkStripCjkAutolink`, before `remarkMath`.
*
* Security note: this plugin deliberately has NO scheme allowlist — it converts
* every `scheme://` pattern it finds. The sanitize layer (`urlTransform`) and
* the onClick layer (`classifyScheme`) handle blocking/confirmation downstream.
* Keeping the plugin allowlist-free means new custom schemes work automatically
* without touching this file.
*/
// Matches a non-http(s) URI of the form `scheme://...` where:
// - scheme is one or more ASCII letters (a-z), digits, +, -, or .
// - scheme is NOT "http" or "https" (those are GFM's domain)
// - followed by "://" and a run of non-whitespace characters
//
// Trailing punctuation (.,!?;:) and closing brackets/parens are stripped
// by a post-match trim step so "See obsidian://x." doesn't include the ".".
const NON_HTTPS_URI_RE = /\b(?!https?:\/\/)([a-zA-Z][a-zA-Z0-9+\-.]*):\/\/[^\s]*/g
// Characters that may be stripped from the end of a matched URI.
// `)` and `]` are only stripped when the URL body has no unmatched opening
// counterpart — this mirrors GFM autolink literal behaviour and keeps URIs
// like `obsidian://open?file=Note(1)` intact.
const TRAILING_PUNCT_CHARS = /[.,;!?:)>\]'"]/
/**
* Strip trailing punctuation from a URI, but preserve `)` / `]` that close an
* unmatched `(` / `[` inside the URL body.
*
* Examples:
* `obsidian://x.` → stripped `obsidian://x`, trailing `.`
* `obsidian://open?file=Note(1)` → stripped unchanged, trailing ``
* `obsidian://x).` → stripped `obsidian://x`, trailing `).` (no `(` to balance)
*/
function stripTrailingPunct(uri: string): { stripped: string; trailing: string } {
let stripped = uri
let trailing = ''
while (stripped.length > 0) {
const last = stripped[stripped.length - 1]
if (!TRAILING_PUNCT_CHARS.test(last)) break
if (last === ')' || last === ']') {
const open = last === ')' ? '(' : '['
const inner = stripped.slice(0, -1)
let opens = 0
let closes = 0
for (const ch of inner) {
if (ch === open) opens++
else if (ch === last) closes++
}
// If the inner URL already has more opens than closes, the trailing
// closer balances an earlier opener and belongs to the URL.
if (closes < opens) break
}
trailing = last + trailing
stripped = stripped.slice(0, -1)
}
return { stripped, trailing }
}
interface MdastNode {
type: string
url?: string
value?: string
lang?: string
children?: MdastNode[]
}
/**
* Walk all text nodes inside paragraph-like containers and replace
* `scheme://...` patterns with link nodes.
*
* Skips:
* - `code` (fenced code blocks) and `inlineCode` nodes — never touched.
* - `link` / `linkReference` nodes — their children are not re-processed
* (existing links are left as-is).
*/
function visitAndLinkify(node: MdastNode): void {
if (!node.children) return
const newChildren: MdastNode[] = []
for (const child of node.children) {
// Don't descend into existing links or code nodes.
if (
child.type === 'link'
|| child.type === 'linkReference'
|| child.type === 'inlineCode'
|| child.type === 'code'
) {
newChildren.push(child)
continue
}
if (child.type === 'text' && typeof child.value === 'string') {
const segments = linkifyText(child.value)
newChildren.push(...segments)
continue
}
// Recurse into other container nodes (e.g. paragraph, blockquote, list items).
visitAndLinkify(child)
newChildren.push(child)
}
node.children = newChildren
}
/**
* Split a raw text string around any `scheme://...` matches and return a
* mixed array of text nodes and link nodes.
*/
function linkifyText(text: string): MdastNode[] {
const result: MdastNode[] = []
let lastIndex = 0
// Reset the regex state (global flag carries state across calls).
NON_HTTPS_URI_RE.lastIndex = 0
let match: RegExpExecArray | null
while ((match = NON_HTTPS_URI_RE.exec(text)) !== null) {
const rawUri = match[0]
const matchStart = match.index
// Strip trailing punctuation characters from the URI, preserving
// balanced ()/[].
const { stripped, trailing } = stripTrailingPunct(rawUri)
// Text before this match.
if (matchStart > lastIndex) {
result.push({ type: 'text', value: text.slice(lastIndex, matchStart) })
}
// The link node.
result.push({
type: 'link',
url: stripped,
children: [{ type: 'text', value: stripped }],
})
// Any stripped trailing punctuation becomes a plain text node.
if (trailing) {
result.push({ type: 'text', value: trailing })
}
lastIndex = matchStart + rawUri.length
}
// Remaining text after the last match.
if (lastIndex < text.length) {
result.push({ type: 'text', value: text.slice(lastIndex) })
}
// If no matches were found, return the original text node unchanged.
if (result.length === 0) {
result.push({ type: 'text', value: text })
}
return result
}
export default function remarkNonHttpsAutolink() {
return (tree: MdastNode) => {
visitAndLinkify(tree)
}
}