Commit Graph
10772 Commits
Author SHA1 Message Date
Aleksander Grygier 57376adf97 feat: UI improvements 2026-09-01 13:44:13 +02:00
Aleksander Grygier 14e836f96e feat: keep sidecar badge as MTP in chips and quant select
Reverts the -shared suffix on the chip badge and drops the shared
prefix from the draft quant select options: both MTP variants read as
plain MTP, distinguished only by their file name tooltip and size.

Assisted-by: Claude Sonnet
2026-09-01 13:33:49 +02:00
Aleksander Grygier 8faff6643c chore: fix lint ordering in models ui components
simple-import-sort import ordering and perfectionist object key order,
from the pre-push hook run.

Assisted-by: Claude Sonnet
2026-09-01 13:15:52 +02:00
Aleksander Grygier 1d7dcda7f9 feat: render capability icons in the meta row of multi-line model ids
With iconsOnNewLine, the tool use / reasoning / modality icons now sit
in the second row, in line with context length and disk size, instead of
competing with the name on the identity line. Icons render through a
shared snippet in both positions; single-line usages keep the icons
inline after the name.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier b5ec1d2dc0 feat: request base_model tags so search rows get base org avatars
getBaseModels reads cardData.base_model / base_model: tags, which the HF
list endpoint omits by default; search rows therefore fell back to the
repo org avatar. Adding tags to the expand list lets search rows render
the base org avatar as main with the quant org as the corner badge, like
catalog rows.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier cb2a9a7f10 feat: mark shared draft variants on the sidecar badge
The shared marker moves from a 'shared' suffix on the quant name to the
sidecar badge (MTP-SHARED); draft select options prefix the badge text
so both variants of the same quant stay distinguishable in the dropdown,
and the download CTA shows e.g. 'Download Q4_K_M + MTP-SHARED'.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 0bbc298d3f fix: restore componentized download options styling
The parent had been overwritten with a pre-componentization version,
dropping the shared quant toggles, the download button/command
components and the section's frosted surface styling. Restores the
componentized version and merges the manual refinements into the child
components: primary-tinted quant selects, taller CTA with press scale,
command box accent bar with an absolutely positioned copy button.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 584b18551b refactor: read the catalog from store state in catalogBuilds
Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier fc682165bc feat: render discover rows from expanded HF list data
Search results now request the fields the rows render (gguf, siblings,
pipeline_tag, ...) via repeated expand params, so a search row shows the
same badges as a catalog row. The catalog path fetches each repo's tree
and derives a per-repo min/max size range (quants plus draft sidecars)
through the models hub store, with search rows falling back to one lazy
tree fetch per repo. Sizes from llama.app catalog size strings parse via
HuggingFaceService.parseSizeBytes when sizeBytes is absent.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 48aeb65026 feat: label shared draft variants in download options
Unsloth ships MTP draft heads in two layouts: shared- files borrow the
embedding/output weights from the target model, others are
self-contained. extractQuantMeta now flags a standalone 'shared'
segment, and the chips, quant selects and download CTA show e.g.
'Q4_K_M shared' next to the plain quant so the two files are
distinguishable.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 95bc47984d feat: single-line scrollable terminal command in download options
Long commands (draft models, long repo ids) no longer wrap to a second
line: the command row scrolls horizontally, tokens never shrink, and the
copy button stays pinned outside the scroll area.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier c3839b01b9 refactor: extract ModelsDiscoverModelDetailsMetadataItem component
One label | value metadata chip (model size, context, architecture,
license) now renders through a dedicated component; the chat template
button and gated badge stay inline in the metadata component as they
are different shapes.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 08f8bd5230 feat: skeleton loading states for model list and details
Replaces the 'Loading model...' text with a static skeleton mirroring
the detail layout (header, metadata chips, download options box, readme
lines). List item skeletons now vary widths per row so the loading list
does not render as identical blocks.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier b72af5b109 fix: detect sidecar GGUFs nested in repo folders
extractQuantMeta parsed the full sibling path, so repo layouts that nest
sidecars in a folder (e.g. MTP/mtp-Model-Q4_0.gguf) failed the sidecar
prefix match and were classified as main weights.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:01:43 +02:00
Aleksander Grygier 7cbbd6ce20 feat: WIP 2026-09-01 11:20:34 +02:00
Aleksander Grygier 1ecfd227ad feat: WIP 2026-09-01 10:34:12 +02:00
Aleksander Grygier cc08bc4b97 chore: Add LLAMA-APP-REUSE comments 2026-09-01 08:24:36 +02:00
Aleksander Grygier 46c1d39b12 feat: Models Downloading UI/UX 2026-09-01 08:15:37 +02:00
Aleksander Grygier 2043d0a903 ui : show downloaded quants and drafts as static chips
Downloaded files are no longer selectable toggles; they render as a
chip with a checkmark instead.

Assisted-by: pi
2026-09-01 01:25:30 +02:00
Aleksander Grygier 173759053c ui : merge download options and terminal commands into one panel
The quant and draft sidecar buttons become a multiple toggle group
(selecting which files to download, not download triggers), and below
the panel the terminal llama serve command updates live to reflect the
selection, with a download CTA that fires the downloads for the
selected entries (main + optional draft sidecar).

Adds the shadcn-svelte toggle-group component.

Assisted-by: pi
2026-09-01 00:48:00 +02:00
Aleksander Grygier 41c67c0c55 ui : add download manager pieces and the list search component
- ModelsDiscoverListSearch: extracted search input from the list
- ModelsDiscoverModelDetailsMetadata: description + metadata chips
  extracted from the details header
- ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors
  embedded in the inline command text
- ModelsDownloadManager: tracked downloads with per-file progress and
  a delete action
- ModelsDownloadManagerDownloadStatusToast: one toast per download
  with a progress bar per file (main + sidecars) and a CTA to open
  the download manager
- DialogModelsDownloadManager: dialog shell for the manager

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier c00a7ccd83 ui : restructure discover components to the target names
- ModelsDiscoverItem + ModelsDiscoverInfo fold into
  ModelsDiscoverListItem (avatar + model id + badges + context/size)
- ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails*
- TerminalCommands renamed to ModelsDiscoverModelDetailsCommands
- ModelsDiscoverDetailsName folded into the details header
- ModelsDiscoverListSearch extracted from the list search input
- stories updated for the new names

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier c43cec9c26 ui : add a Discover models trigger to the model selector
Add a Discover models item to the selector dropdown footer that opens
the DialogModelsDiscover dialog.

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier 8bf37609b0 ui : wire the discover dialog to live data and downloads
Load the selected model's details, file tree and README via
HuggingFaceService on selection change, and expose the download
progress type globally for the status feed. The download options
already read their state from the models status store.

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier 71a379277b ui : simplify download options to plain memory estimate
Replace the device-memory tier badges with the simple memory estimate:
each quant tooltip shows the estimated runtime memory and the device/OS
chip is dropped, matching the estimateModelMemoryBytes model. Make the
download dialog callbacks optional so the component stays presentational
until wired to the live status store.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier b8f0ef583d ui : add models discover components with stories
Port the discover UI from the scrapbook, adapted to the typed
sidecar API: searchable two-pane explorer (list, item, info, org
avatar with quant badge), model details (header, name badges,
download options grouped by bit depth with compatibility tiers,
terminal serve/cli commands per draft sidecar, README viewer, chat
template dialog), download confirmation dialog with progress, and
the full-screen dialog shell.

Presentational components take data and download state via props;
the loading container and store wiring land in the integration
branch. MarkdownContent gains a sanitized allowHtml option used by
the README viewer; ModelId gains context, size-range, params and
sidecar badges; the selector option passes thinking/tool flags
instead of the removed capabilities prop.

Basic Storybook stories cover each component with HF-shaped
fixtures.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 70f8d049f6 ui : mark llama-app-reusable code with a LLAMA-APP-REUSE tag
Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 5c64869570 ui : mark llama-app-reusable code with a LLAMA-APP-REUSE tag
Tag the pure-logic files and functions that llama.app (llama-pages)
can reuse as-is: model id parsing, HF name and quant conventions,
hardware compatibility estimation, chat-template capability
detectors, and the HF formatting and metadata helpers. App-specific
code is left unmarked.

The LLAMA-APP-REUSE prefix makes the reusable surface greppable and
distinguishable from regular comments: grep -rn LLAMA-APP-REUSE.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 5b7a4340f7 ui : fix double blank line in model types
Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 14ac965b2e ui : add download state tracking to the model status manager
Port the download lifecycle into ModelsStatusManager: track per-entry
progress keyed by <repo>:<tag> from the /models/sse feed, record
failed downloads for the delete-and-retry path, and expose the
downloadModel / cancelDownload operations (POST/DELETE /models).

Add the ServerModelStatus.DOWNLOADED/DOWNLOADING cases and the
ModelDownloadProgress type.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 0d22938cd9 ui : add model download pipeline
Wire the model download flow: ModelsService.downloadModel (POST
/models) and cancelDownload (DELETE /models), the apiDelete helper,
ApiModelsDownloadRequest/Response types, the download_progress SSE
payload, and the download_finished/download_failed SSE event kinds
matching the server feed.

Add modelsHubStore owning the HuggingFace GGUF model list for the
discover dialog: curated catalog defaults on open, search replaces
the list across all of HuggingFace.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier a07359ef93 ui : simplify model memory estimation
Replace the device-memory tier machinery with a plain file-size
estimate: required runtime memory is the model file size with
headroom for KV cache and allocator overhead (estimateModelMemoryBytes).
Callers present the requirement; there is no device detection and no
fit-versus-budget verdict.

Drops resolveDeviceMemoryGb, deviceMemoryBudgetMb,
computeFileCompatibilityTiers and the CompatibilityTier type, and the
barrel keeps only the new estimator.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier af98525e4c ui : add model compatibility estimation
Port the hardware-compatibility estimator from ggml-org/llama-macos:
map every GGUF file in a repo to a full/limited/none tier based on
the device memory budget (GPU working set approximated from RAM, less
fit slack and an OS floor) and the estimated weight + context memory.

Main quants are tiered individually; shards, mmproj and quant-matched
draft sidecars inherit their main quant's tier. Sidecar picking
mirrors the server's find_best_sibling ranking (deepest directory,
exact quant tag, closest bit depth).

Also port detectToolUseSupport (infers tool-calling support from a
chat template) and the browser get_info fallback helper.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 548acc2ffd ui : move catalog url to models-discover constants
Address review follow-up: the llama.app catalog endpoint belongs to
the models-discover feature, not the HF constants. Use Number() for
shard index parsing and name the UD-quant prefix segment lookup.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 7e63e9e346 ui : move HF constants to constants and enums modules
Address review on the HF data layer:

- replace the HfModelSort / SidecarForm / sibling entry type string
  unions with enums (HfModelSort, SidecarForm, HfEntryType)
- move URLs, query params, regexes, limits, retry settings, shard
  file conventions, tag tokens and formatting units into a dedicated
  huggingface.constants.ts; reuse the existing PATH_SEPARATOR
- drop the task label / pipeline icon / library display maps: the
  discover UI only presents GGUF models, so keep the task tags for
  logic use only (parseTags)
- drop the hardcoded curated model list; the discover dialog gets its
  default list from the llama.app /v1/catalog.json endpoint, which is
  an acceptable online-only source since the feature requires internet
  access anyway

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier ddb8d219e1 ui : follow MODEL_ID regex rename in huggingface service
Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 4aa7063be3 ui : add huggingface hub data layer
Add HuggingFaceService for browsing and searching GGUF models on the
HF Hub: catalog/model search, model details, repo file tree, raw
README fetch, and the llama.app model catalog. Includes GGUF file
analysis helpers - extractQuantMeta (quant token plus sidecar type
and its form, prefix or suffix), shard collapsing, quant bit-depth
lookup, and download/size/likes formatting.

Add the HF API types and the curated model list shown in the
Discover Models sidebar.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier f6f90463be ui : lowercase reasoning capability and use sidecar enums in tests
Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 4c93cef4f5 ui : address review on model id parsing
Use lowercase values for the sidecar enums so the value doubles as
the filename token, derive the sidecar regexes from the enum values,
and rename the MODEL_ID regex keys to the _REGEX suffix used by the
rest of the constants files. Replace the tools capability magic
string with ModelCapability.TOOL_USE.

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier c1cbd5e277 ui : parse sidecar types in model ids
Add ModelDraftSidecar / ModelAuxSidecar enums with a ModelSidecar
union type; mmproj is the only auxiliary sidecar (single member,
covers vision and audio input). Add SIDECAR_PREFIX/SUFFIX_RE regex
matching the server's filename conventions, and type guards +
enum-file-token helpers in model-id.constants.ts.

Extend parseModelId to detect sidecar filename tokens (mtp-, mmproj-,
etc) and expose isDraftSidecar / isAuxSidecar / sidecarFromFileToken
helpers. Add ModelCapability.TOOL_USE with icon/label/flag mappings.

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 4b5449a445 ui : add format:files script for formatting changed files only
Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 9a8519788b ui : rename model API types to mirror endpoints
Drop the Router prefix from client-side API types; names now map
directly to the /models endpoint family (load/unload/download/list).
Merge ApiModelListResponse into ApiModelsListResponse (same endpoint
shape in both modes) and remove the duplicate ModelsService.listRouter().

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 2853c1816c ui : remove dead ChatFormActionAddMcpSubmenu component
Unreferenced leftover from the pre-dialog MCP design; references
context props that no longer exist and breaks svelte-check.

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Buğra Özgürsoy 458681e1d5 metal : add fa-vec tunings for M1 Ultra (#28088)
* metal : add fa-vec tunings for M1 Ultra

* metal : move M1 Ultra tunings after M1 Max section

* metal : remove duplicate blank line
b10729
2026-08-31 23:47:27 +02:00
ynankani e4b9af007b CUDA: XOR swizzle flash attn K,V smem fp16 tiles (#25635)
* CUDA: XOR swizzle flash attn  K,V smem fp16 tiles

Signed-off-by: ynankani <[email protected]>

* Fix use 64bit generic pointer instead of 32bit shared pointer

Signed-off-by: ynankani <[email protected]>

* fix shared memory race in FA on DGX Spark

* Handle corener case

Signed-off-by: ynankani <[email protected]>

* Add swizzle test cases and gate sync for swizzled path only

Signed-off-by: ynankani <[email protected]>

* gate CUDA PTX

Signed-off-by: ynankani <[email protected]>

* offset calculation specific for swizzle branch

Signed-off-by: ynankani <[email protected]>

* Reafctor code

Signed-off-by: ynankani <[email protected]>

* Refactor FA swizzle ldmatrix if/else into helpers (K row/col, V offset)

Signed-off-by: ynankani <[email protected]>

* rebase and update test case args

Signed-off-by: ynankani <[email protected]>

* Allow swizzle for non-pow2 shapes, for which nbatch_2%32==0

Signed-off-by: ynankani <[email protected]>

---------

Signed-off-by: ynankani <[email protected]>
b10728
2026-08-31 22:18:01 +02:00
Georgi Gerganov ab0b3bd3c8 metal : add concat support for quantized types (#28116)
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-0731
b10727
2026-08-31 23:16:04 +03:00
BartowskiandGeorgi Gerganov 85c55223ca AVX2: Speed up large batch size prompt processing of IQ models (#27402)
* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <[email protected]>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991e.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <[email protected]>
b10726
2026-08-31 14:33:50 -04:00
Georgi Gerganov 2a74817f93 metal : add top-k radix implementation (#28073)
Assisted-by: DeepSeek-v4-Flash-0731
2026-08-31 21:31:53 +03:00
itsnotoger 2d8d612e4c kv-cache : optimize restoring non-contiguous cells (#27991)
* kv cache : batch state restore scatter reads per contiguous run

When restoring state into non-contiguous destination cells (e.g. a
prompt-cache snapshot into a fragmented ring), state_read_data issued
one small copy per KV cell - ~1.4M copies of a few KiB each for a
40k+ token restore, taking 25-63 s on the CUDA backend.

The snapshot stores cell rows in cell order, so a maximal run of
consecutive destination indices maps to one contiguous block and can
be restored with a single copy. Precompute the runs once and use them
in all three scatter loops (K, V, transposed V). Byte-identical.

The on-device reader copies with a byte cursor when the read and
write chunking differs, so the batched reads are safe for it as well.
Batching makes equal tensor counts with a different split reachable
(save ranges [2,1] vs restore runs [1,2]); the next commit teaches the
reader's 1:1 path to fall back to the byte cursor in that case.

Verified in a production setup: 1,363,616 copies / 25-63 s -> 224
copies / 221-424 ms for the same restores (42,603 cells, 4 runs).

Assisted-by: Claude Code (unsloth/qwen3.8-27b)

* context : fall back to the byte cursor when read and write chunking differ

the on-device reader copies saved state back with a 1:1 copy by tensor
index whenever the write and read sides recorded the same number of
tensors, guarded by a per-tensor size assert.

equal tensor counts do not imply equal chunking: a state restore may
batch its reads per contiguous run of destination cells while the save
used per-range reads, so both sides can record two tensors that split
the same data differently, and the assert aborts in all builds.

compare the per-tensor sizes and only take the 1:1 path when the
chunking actually matches, otherwise fall through to the existing
byte-cursor copy. both sides enumerate the same logical data in the
same order, so the cursor copy is well-defined across tensor
boundaries.

Assisted-by: Claude Code (unsloth/qwen3.8-27b)

* tests : cover state restore scatter reads on host and on-device paths

decode the same prefix on two sequences, interleaving the seq 0 cells
between the seq 1 cells, so the seq 1 cells are isolated from each
other in the kv cache (three cells, two saved ranges). save the seq 1
state, free the interleaved seq 0 cells, and restore: the destination
is then non-contiguous (two runs), and the restore-side chunking has
the same tensor count as the save-side with a different split, so the
scatter path is batched per contiguous run and the on-device reader's
byte-cursor fallback is exercised.

the restored state is saved again on the host and compared byte for
byte with the first save: the blob is serialized in sequence cell
order, so the two saves are identical if and only if the scatter
restore wrote exactly the same KV content. this documents the
byte-identical guarantee of the run-batched scatter reads.

one test per io backend: the host (CPU) path and the on-device path.

Assisted-by: Claude Code (unsloth/qwen3.8-27b)
b10724
2026-08-31 19:49:58 +03:00
Hongqiang Wang 010be9683a opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (#26438)
* opencl: Q4_K/Q5_K mul_mv N_DST 4->8 on Intel for 2x activation reuse

* opencl: Q4_K mul_mm 8x8 tile fot Intel

* opencl: Q5_K mul_mm 8x8 tile for Intel

* opencl: Q4_K mul_mv N_DST 8->16 for Intel
b10723
2026-08-31 08:56:22 -07:00