Files
plainleaf/web/markdown_parser/parser.ts
T
aa0b95b31d The one about refs (#1496)
* [Major] Introduce a new `Ref` and `Path` type

This introduces a new `Ref` and `Path` type, which try to be as lightweight and rigorous as possible. More description in the PR

* [Minor] Updated Syntax

* [Minor] Updated Type

* [Minor] Renamed file

* [Minor] Updated type

* [Minor] Updated Syntax

* [Major] Rename css classes for wiki links

Previously all the parts of a wiki link were tagged with `sb-wiki-link-page`. This didn't make a whole lot of sense. Also the tag `sb-wiki-link` was could never render, because of the way `HighlightStyle` works. Now every part of a wiki link will be tagged with `sb-wiki-link` (i.e. the link, the alias, the dimensions). Invalid and missing links will be styled separately in a later commit.

* [Minor] Adjust to new API

* [Minor] Small changes

* [Minor] Function and file renames

* [Major] Syscall adjusments

Change behavior of `getCurrentPath` to actually return what's now considered a `Path`. `getCurrentPage(Meta)` now also correctly returns for documents. We could consider a rename to get rid of the `page` association

* [Major] Rework `Client`

Mostly minor adjustments for the new API. Some major changes in `navigateWithinPage` and some unecessary checks removed now

* [Major] Errors on navigate

This now throws an error when the user tries to navigate to a link that is invalid

* [Major] Minor adjustments and remove some old bulk

This does the typicall adjustments to the new API. A thing to consider is that invalid links aren't indexed. Also removed old code about template and query stuff. And reused a regex instead of copy pasting it.

* [Minor] Cleanup editor ui

* [Minor] Adjust to new API

* [Minor] Adjust to new API and reduce ugly regex

* [Minor] New API

* [Minor] Rework link cm plugin

* [Major] Rework wiki links

A lot of code cleanup, reworked for new APIs and also introduces "invalid links"

* [Minor] Part of the editor ui rework

* [Major] Adjust to new API

This also removes the `looksLikePathWithExtension` check. I don't think it was necessary.

* [Minor] PArt of editor ui rewrite

* [Minor] Adjust to new API

* [Major] Adjust renderer to new API

Invalid links now link to `#`

* [Major] Rewrite navigator

This still needs major rework, currently in a "got it working" state

* [Minor] Adjust file name

* [Minor] Mistake during rebase

* [Minor] Navigate to end of header

* [Minor] Remove `tweakEditorDOM` and use proper CM extensions

* [Major] Heavily clean up the navigation code

* [Minor] Make `PageNavigator` closable again

* [Minor] Small fixes to reduce errors bubbling up

* [Minor] Show error message if navigate fails

* [Major] Remove legacy `navigate({ kind = ` syntax

This removes the old navigation syntax. It's still supported but the user will be warned then using it. I also added some validation to editor.navigate for user provided refs.

* [Minor] Make notifications errors

* [Minor] Don't fail if header isn't found

* [Minor] Typo

* [Major] Remove cleanPageRef

* [Minor] Small bug

* [Minor] Align behavior between render and editor

* [Major] Remove `PageRef`. Pretty sure it's obsolete

* [Major] Remove pos from header as it's an invalid ref

* [Minor] Align with original behavior

* [Minor] Get tests working again

* [Minor] check that indexPage is valid

* [Major] Rewrite `resolve.ts` file

* [Minor] More tests, more fixes

* [Minor] Don't just append .md this is weird

* [Minor] More edge cases

* [Minor] Version bumping hono somehow fixes the `/` case

* [Minor] Smaller TODOs

* [Major] We can be a littler looser on the restrictions

* [Major] Change header indexing, to index with pos

* [Minor] More fixes for headers

* [Minor] Rename for clarity

* [Minor] Change more regexes to named groups

* [Major] Inline content rewrite

* [Minor] Fix TODO

* [Minor] Bump index version

* [Major] New Space Primitive to check paths

* [Minor] Fix #1214

* [Minor] Fix #1169

* [Minor] Remove TODO

This check is probably not needed anymore. Unsure. May look into it later

* [Minor] small bug

* [Minor] Not a good rebase without me needing to fix it afterwards

* [Minor] Improve fallback for file upload

* [Minor] Some fixes on the for resolving

* [Minor] Removed doubled test case

* [Minor] Resolve #860

* [Minor] Turn flash notification into an error

* [Minor] Change handling of `.` paths

* [Minor] Show empty URL for index page

* [Minor] Push correct ref to history

* [Major] Capture mini editor events in capture phase

This fixes the issue were when pressing arrowup or arrowdown changed the cursor position inside the editor

* [Major] Add prefix to `FilterList`

This now fixes issues were if you had a prefixed page you  con't properly search for it, because the search considered the prefix as part of the name

* [Minor] A little rigor for empty headers

* [Major] Fix #1404

* [Minor] `parseInt` returns `NaN` handle that

* [Minor] Use `fileName`

* [Minor] Fix comment

* [Minor] Fixes

* [Major] disallow `foo.png.md` and `/../`, etc. cases

* [Major] Change name of conflicted pages

* [Minor] Remove unecessary check

* [Minor] Remove some split code

* [Minor] Allow deletion of invalid files

* [Minor] Bug in the regex

* [Minor] Give some possibility of fixing incorrect names

* [Minor] Tiny bug with header completion

* [Major] Docs

* [Minor] Ignore errors for navigation inside of a page

This could be a problem when SB e.g. cached a scrollTop for a page, which, because e.g. the file is now deleted, doesn't exist anymore and thus CM throws an error, when trying to navigate

* Update conflicted copy regex in Maintenance.md

---------

Co-authored-by: Zef Hemel <zef@zef.me>
2025-08-24 16:16:19 +02:00

438 lines
12 KiB
TypeScript

import { yaml as yamlLanguage } from "@codemirror/legacy-modes/mode/yaml";
import { styleTags, type Tag, tags as t } from "@lezer/highlight";
import {
type Line,
type MarkdownConfig,
Strikethrough,
Subscript,
Superscript,
} from "@lezer/markdown";
import { markdown } from "@codemirror/lang-markdown";
import { foldNodeProp, StreamLanguage } from "@codemirror/language";
import * as ct from "./customtags.ts";
import { NakedURLTag } from "./customtags.ts";
import { TaskList } from "./extended_task.ts";
import { Table } from "./table_parser.ts";
import { pWikiLinkRegex, tagRegex } from "./constants.ts";
import { parse } from "./parse_tree.ts";
import type { ParseTree } from "@silverbulletmd/silverbullet/lib/tree";
import { luaLanguage } from "../../lib/space_lua/parse.ts";
const WikiLink: MarkdownConfig = {
defineNodes: [
{ name: "WikiLink" },
{ name: "WikiLinkPage", style: ct.WikiLinkPartTag },
{ name: "WikiLinkAlias", style: ct.WikiLinkPartTag },
{ name: "WikiLinkDimensions", style: ct.WikiLinkPartTag },
{ name: "WikiLinkMark", style: t.processingInstruction },
],
parseInline: [
{
name: "WikiLink",
parse(cx, next, pos) {
// Do a preliminary check for performance
if (next != 91 /* '[' */ && next != 33 /* '!' */) {
return -1;
}
pWikiLinkRegex.lastIndex = 0;
const match = pWikiLinkRegex.exec(cx.slice(pos, cx.end));
if (!match || !match.groups) {
return -1;
}
//const [fullMatch, firstMark, page, alias, _lastMark] = match;
const { leadingTrivia, stringRef, alias } = match.groups;
const endPos = pos + match[0].length;
let aliasElts: any[] = [];
if (alias) {
const pipeStartPos = pos + leadingTrivia.length + stringRef.length;
aliasElts = [
cx.elt("WikiLinkMark", pipeStartPos, pipeStartPos + 1),
cx.elt(
"WikiLinkAlias",
pipeStartPos + 1,
pipeStartPos + 1 + alias.length,
),
];
}
let allElts = cx.elt("WikiLink", pos, endPos, [
cx.elt("WikiLinkMark", pos, pos + leadingTrivia.length),
cx.elt(
"WikiLinkPage",
pos + leadingTrivia.length,
pos + leadingTrivia.length + stringRef.length,
),
...aliasElts,
cx.elt("WikiLinkMark", endPos - 2, endPos),
]);
// If inline image
if (next == 33) {
allElts = cx.elt("Image", pos, endPos, [allElts]);
}
return cx.addElement(allElts);
},
after: "Emphasis",
},
],
};
const LuaDirectives: MarkdownConfig = {
defineNodes: [
{ name: "LuaDirective" },
{ name: "LuaExpressionDirective" },
{ name: "LuaDirectiveMark", style: ct.DirectiveMarkTag },
],
parseInline: [
{
name: "LuaDirective",
parse(cx, next, pos) {
const textFromPos = cx.slice(pos, cx.end);
if (
next !== 36 /* '$' */ ||
cx.slice(pos, pos + 2) !== "${"
) {
return -1;
}
let bracketNestingDepth = 0;
let valueLength = 0;
// We need to ensure balanced { and } pairs
loopLabel:
for (; valueLength < textFromPos.length; valueLength++) {
switch (textFromPos[valueLength]) {
case "{":
bracketNestingDepth++;
break;
case "}":
bracketNestingDepth--;
if (bracketNestingDepth === 0) {
// Done!
break loopLabel;
}
break;
}
}
if (bracketNestingDepth !== 0) {
return -1;
}
const bodyText = textFromPos.slice(2, valueLength);
const endPos = pos + valueLength + 1;
// Let's parse as an expression
const parsedExpression = luaLanguage.parser.parse(`_(${bodyText})`);
// If bodyText starts with whitespace, we need to offset this later
const whiteSpaceOffset = bodyText.match(/^\s*/)?.[0].length ?? 0;
const node = parsedExpression.resolveInner(2, 0).firstChild?.nextSibling
?.nextSibling;
if (!node) {
return -1;
}
const bodyEl = cx.elt(
"LuaExpressionDirective",
pos + 2,
endPos - 1,
[cx.elt(node.toTree()!, pos + 2 + whiteSpaceOffset)],
);
return cx.addElement(
cx.elt("LuaDirective", pos, endPos, [
cx.elt("LuaDirectiveMark", pos, pos + 2),
bodyEl,
cx.elt("LuaDirectiveMark", endPos - 1, endPos),
]),
);
},
after: "Emphasis",
},
],
};
const HighlightDelim = { resolve: "Highlight", mark: "HighlightMark" };
export const Highlight: MarkdownConfig = {
defineNodes: [
{
name: "Highlight",
style: { "Highlight/...": ct.Highlight },
},
{
name: "HighlightMark",
style: t.processingInstruction,
},
],
parseInline: [
{
name: "Highlight",
parse(cx, next, pos) {
if (next != 61 /* '=' */ || cx.char(pos + 1) != 61) return -1;
return cx.addDelimiter(HighlightDelim, pos, pos + 2, true, true);
},
after: "Emphasis",
},
],
};
export const attributeStartRegex = /^\[([\w\$]+)(::?\s*)/;
export const Attribute: MarkdownConfig = {
defineNodes: [
{ name: "Attribute", style: { "Attribute/...": ct.AttributeTag } },
{ name: "AttributeName", style: ct.AttributeNameTag },
{ name: "AttributeValue", style: ct.AttributeValueTag },
{ name: "AttributeMark", style: t.processingInstruction },
{ name: "AttributeColon", style: t.processingInstruction },
],
parseInline: [
{
name: "Attribute",
parse(cx, next, pos) {
let match: RegExpMatchArray | null;
const textFromPos = cx.slice(pos, cx.end);
if (
next != 91 /* '[' */ ||
// and match the whole thing
!(match = attributeStartRegex.exec(textFromPos))
) {
return -1;
}
const [fullMatch, attributeName, attributeColon] = match;
let bracketNestingDepth = 1;
let valueLength = fullMatch.length;
loopLabel:
for (; valueLength < textFromPos.length; valueLength++) {
switch (textFromPos[valueLength]) {
case "[":
bracketNestingDepth++;
break;
case "]":
bracketNestingDepth--;
if (bracketNestingDepth === 0) {
// Done!
break loopLabel;
}
break;
}
}
if (bracketNestingDepth !== 0) {
console.log("Failed to parse attribute", fullMatch, textFromPos);
return -1;
}
if (textFromPos[valueLength + 1] === "(") {
// This turns out to be a link, back out!
return -1;
}
return cx.addElement(
cx.elt("Attribute", pos, pos + valueLength + 1, [
cx.elt("AttributeMark", pos, pos + 1), // [
cx.elt("AttributeName", pos + 1, pos + 1 + attributeName.length),
cx.elt(
"AttributeColon",
pos + 1 + attributeName.length,
pos + 1 + attributeName.length + attributeColon.length,
),
cx.elt(
"AttributeValue",
pos + 1 + attributeName.length + attributeColon.length,
pos + valueLength,
),
cx.elt("AttributeMark", pos + valueLength, pos + valueLength + 1), // [
]),
);
},
after: "Emphasis",
},
],
};
type RegexParserExtension = {
// unicode char code for efficiency .charCodeAt(0)
firstCharCode: number;
regex: RegExp;
nodeType: string;
tag: Tag;
className?: string;
};
function regexParser({
regex,
firstCharCode,
nodeType,
}: RegexParserExtension): MarkdownConfig {
return {
defineNodes: [nodeType],
parseInline: [
{
name: nodeType,
parse(cx, next, pos) {
if (firstCharCode !== next) {
return -1;
}
const match = regex.exec(cx.slice(pos, cx.end));
if (!match) {
return -1;
}
return cx.addElement(cx.elt(nodeType, pos, pos + match[0].length));
},
},
],
};
}
const NakedURL = regexParser(
{
firstCharCode: 104, // h
regex:
/(^https?:\/\/([-a-zA-Z0-9@:%_\+~#=]|(?:[.](?!(\s|$)))){1,256})(([-a-zA-Z0-9(@:%_\+~#?&=\/]|(?:[.,:;)](?!(\s|$))))*)/,
nodeType: "NakedURL",
className: "sb-naked-url",
tag: NakedURLTag,
},
);
const Hashtag = regexParser({
firstCharCode: 35, // #
regex: new RegExp(`^${tagRegex.source}`),
nodeType: "Hashtag",
className: "sb-hashtag-text",
tag: ct.HashtagTag,
});
const TaskDeadline = regexParser({
firstCharCode: 55357, // 📅
regex: /^📅\s*\d{4}\-\d{2}\-\d{2}/,
className: "sb-task-deadline",
nodeType: "DeadlineDate",
tag: ct.TaskDeadlineTag,
});
// FrontMatter parser
const yamlLang = StreamLanguage.define(yamlLanguage);
export const FrontMatter: MarkdownConfig = {
defineNodes: [
{ name: "FrontMatter", block: true },
{ name: "FrontMatterMarker" },
{ name: "FrontMatterCode" },
],
parseBlock: [{
name: "FrontMatter",
parse: (cx, line: Line) => {
if (cx.parsedPos !== 0) {
return false;
}
if (line.text !== "---") {
return false;
}
const frontStart = cx.parsedPos;
const elts = [
cx.elt(
"FrontMatterMarker",
cx.parsedPos,
cx.parsedPos + line.text.length + 1,
),
];
cx.nextLine();
const startPos = cx.parsedPos;
let endPos = startPos;
let text = "";
let lastPos = cx.parsedPos;
do {
text += line.text + "\n";
endPos += line.text.length + 1;
cx.nextLine();
if (cx.parsedPos === lastPos) {
// End of file, no progress made, there may be a better way to do this but :shrug:
return false;
}
lastPos = cx.parsedPos;
} while (line.text !== "---");
const yamlTree = yamlLang.parser.parse(text);
elts.push(
cx.elt("FrontMatterCode", startPos, endPos, [
cx.elt(yamlTree, startPos),
]),
);
endPos = cx.parsedPos + line.text.length;
elts.push(cx.elt(
"FrontMatterMarker",
cx.parsedPos,
cx.parsedPos + line.text.length,
));
cx.nextLine();
cx.addElement(cx.elt("FrontMatter", frontStart, endPos, elts));
return true;
},
before: "HorizontalRule",
}],
};
export const extendedMarkdownLanguage = markdown({
extensions: [
WikiLink,
Attribute,
FrontMatter,
TaskList,
Highlight,
LuaDirectives,
Strikethrough,
Table,
NakedURL,
Hashtag,
TaskDeadline,
Superscript,
Subscript,
{
props: [
foldNodeProp.add({
// Don't fold at the list level
BulletList: () => null,
OrderedList: () => null,
// Fold list items
ListItem: (tree, state) => ({
from: state.doc.lineAt(tree.from).to,
to: tree.to,
}),
// Fold frontmatter
FrontMatter: (tree) => ({
from: tree.from,
to: tree.to,
}),
}),
styleTags({
Task: ct.TaskTag,
TaskMark: ct.TaskMarkTag,
Comment: ct.CommentTag,
"Subscript": ct.SubscriptTag,
"Superscript": ct.SuperscriptTag,
"TableDelimiter StrikethroughMark": t.processingInstruction,
"TableHeader/...": t.heading,
TableCell: t.content,
CodeInfo: ct.CodeInfoTag,
HorizontalRule: ct.HorizontalRuleTag,
Hashtag: ct.HashtagTag,
NakedURL: ct.NakedURLTag,
DeadlineDate: ct.TaskDeadlineTag,
NamedAnchor: ct.NamedAnchorTag,
}),
],
},
],
}).language;
export function parseMarkdown(text: string): ParseTree {
return parse(extendedMarkdownLanguage, text);
}