Commit Graph
10764 Commits
Author SHA1 Message Date
Aleksander Grygier fc682165bc feat: render discover rows from expanded HF list data
Search results now request the fields the rows render (gguf, siblings,
pipeline_tag, ...) via repeated expand params, so a search row shows the
same badges as a catalog row. The catalog path fetches each repo's tree
and derives a per-repo min/max size range (quants plus draft sidecars)
through the models hub store, with search rows falling back to one lazy
tree fetch per repo. Sizes from llama.app catalog size strings parse via
HuggingFaceService.parseSizeBytes when sizeBytes is absent.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 48aeb65026 feat: label shared draft variants in download options
Unsloth ships MTP draft heads in two layouts: shared- files borrow the
embedding/output weights from the target model, others are
self-contained. extractQuantMeta now flags a standalone 'shared'
segment, and the chips, quant selects and download CTA show e.g.
'Q4_K_M shared' next to the plain quant so the two files are
distinguishable.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 95bc47984d feat: single-line scrollable terminal command in download options
Long commands (draft models, long repo ids) no longer wrap to a second
line: the command row scrolls horizontally, tokens never shrink, and the
copy button stays pinned outside the scroll area.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier c3839b01b9 refactor: extract ModelsDiscoverModelDetailsMetadataItem component
One label | value metadata chip (model size, context, architecture,
license) now renders through a dedicated component; the chat template
button and gated badge stay inline in the metadata component as they
are different shapes.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier 08f8bd5230 feat: skeleton loading states for model list and details
Replaces the 'Loading model...' text with a static skeleton mirroring
the detail layout (header, metadata chips, download options box, readme
lines). List item skeletons now vary widths per row so the loading list
does not render as identical blocks.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:02:41 +02:00
Aleksander Grygier b72af5b109 fix: detect sidecar GGUFs nested in repo folders
extractQuantMeta parsed the full sibling path, so repo layouts that nest
sidecars in a folder (e.g. MTP/mtp-Model-Q4_0.gguf) failed the sidecar
prefix match and were classified as main weights.

Assisted-by: pi:GLM-5.3-Flash
2026-09-01 13:01:43 +02:00
Aleksander Grygier 7cbbd6ce20 feat: WIP 2026-09-01 11:20:34 +02:00
Aleksander Grygier 1ecfd227ad feat: WIP 2026-09-01 10:34:12 +02:00
Aleksander Grygier cc08bc4b97 chore: Add LLAMA-APP-REUSE comments 2026-09-01 08:24:36 +02:00
Aleksander Grygier 46c1d39b12 feat: Models Downloading UI/UX 2026-09-01 08:15:37 +02:00
Aleksander Grygier 2043d0a903 ui : show downloaded quants and drafts as static chips
Downloaded files are no longer selectable toggles; they render as a
chip with a checkmark instead.

Assisted-by: pi
2026-09-01 01:25:30 +02:00
Aleksander Grygier 173759053c ui : merge download options and terminal commands into one panel
The quant and draft sidecar buttons become a multiple toggle group
(selecting which files to download, not download triggers), and below
the panel the terminal llama serve command updates live to reflect the
selection, with a download CTA that fires the downloads for the
selected entries (main + optional draft sidecar).

Adds the shadcn-svelte toggle-group component.

Assisted-by: pi
2026-09-01 00:48:00 +02:00
Aleksander Grygier 41c67c0c55 ui : add download manager pieces and the list search component
- ModelsDiscoverListSearch: extracted search input from the list
- ModelsDiscoverModelDetailsMetadata: description + metadata chips
  extracted from the details header
- ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors
  embedded in the inline command text
- ModelsDownloadManager: tracked downloads with per-file progress and
  a delete action
- ModelsDownloadManagerDownloadStatusToast: one toast per download
  with a progress bar per file (main + sidecars) and a CTA to open
  the download manager
- DialogModelsDownloadManager: dialog shell for the manager

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier c00a7ccd83 ui : restructure discover components to the target names
- ModelsDiscoverItem + ModelsDiscoverInfo fold into
  ModelsDiscoverListItem (avatar + model id + badges + context/size)
- ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails*
- TerminalCommands renamed to ModelsDiscoverModelDetailsCommands
- ModelsDiscoverDetailsName folded into the details header
- ModelsDiscoverListSearch extracted from the list search input
- stories updated for the new names

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier c43cec9c26 ui : add a Discover models trigger to the model selector
Add a Discover models item to the selector dropdown footer that opens
the DialogModelsDiscover dialog.

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier 8bf37609b0 ui : wire the discover dialog to live data and downloads
Load the selected model's details, file tree and README via
HuggingFaceService on selection change, and expose the download
progress type globally for the status feed. The download options
already read their state from the models status store.

Assisted-by: pi
2026-09-01 00:22:58 +02:00
Aleksander Grygier 71a379277b ui : simplify download options to plain memory estimate
Replace the device-memory tier badges with the simple memory estimate:
each quant tooltip shows the estimated runtime memory and the device/OS
chip is dropped, matching the estimateModelMemoryBytes model. Make the
download dialog callbacks optional so the component stays presentational
until wired to the live status store.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier b8f0ef583d ui : add models discover components with stories
Port the discover UI from the scrapbook, adapted to the typed
sidecar API: searchable two-pane explorer (list, item, info, org
avatar with quant badge), model details (header, name badges,
download options grouped by bit depth with compatibility tiers,
terminal serve/cli commands per draft sidecar, README viewer, chat
template dialog), download confirmation dialog with progress, and
the full-screen dialog shell.

Presentational components take data and download state via props;
the loading container and store wiring land in the integration
branch. MarkdownContent gains a sanitized allowHtml option used by
the README viewer; ModelId gains context, size-range, params and
sidecar badges; the selector option passes thinking/tool flags
instead of the removed capabilities prop.

Basic Storybook stories cover each component with HF-shaped
fixtures.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 70f8d049f6 ui : mark llama-app-reusable code with a LLAMA-APP-REUSE tag
Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 5c64869570 ui : mark llama-app-reusable code with a LLAMA-APP-REUSE tag
Tag the pure-logic files and functions that llama.app (llama-pages)
can reuse as-is: model id parsing, HF name and quant conventions,
hardware compatibility estimation, chat-template capability
detectors, and the HF formatting and metadata helpers. App-specific
code is left unmarked.

The LLAMA-APP-REUSE prefix makes the reusable surface greppable and
distinguishable from regular comments: grep -rn LLAMA-APP-REUSE.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 5b7a4340f7 ui : fix double blank line in model types
Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 14ac965b2e ui : add download state tracking to the model status manager
Port the download lifecycle into ModelsStatusManager: track per-entry
progress keyed by <repo>:<tag> from the /models/sse feed, record
failed downloads for the delete-and-retry path, and expose the
downloadModel / cancelDownload operations (POST/DELETE /models).

Add the ServerModelStatus.DOWNLOADED/DOWNLOADING cases and the
ModelDownloadProgress type.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 0d22938cd9 ui : add model download pipeline
Wire the model download flow: ModelsService.downloadModel (POST
/models) and cancelDownload (DELETE /models), the apiDelete helper,
ApiModelsDownloadRequest/Response types, the download_progress SSE
payload, and the download_finished/download_failed SSE event kinds
matching the server feed.

Add modelsHubStore owning the HuggingFace GGUF model list for the
discover dialog: curated catalog defaults on open, search replaces
the list across all of HuggingFace.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier a07359ef93 ui : simplify model memory estimation
Replace the device-memory tier machinery with a plain file-size
estimate: required runtime memory is the model file size with
headroom for KV cache and allocator overhead (estimateModelMemoryBytes).
Callers present the requirement; there is no device detection and no
fit-versus-budget verdict.

Drops resolveDeviceMemoryGb, deviceMemoryBudgetMb,
computeFileCompatibilityTiers and the CompatibilityTier type, and the
barrel keeps only the new estimator.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier af98525e4c ui : add model compatibility estimation
Port the hardware-compatibility estimator from ggml-org/llama-macos:
map every GGUF file in a repo to a full/limited/none tier based on
the device memory budget (GPU working set approximated from RAM, less
fit slack and an OS floor) and the estimated weight + context memory.

Main quants are tiered individually; shards, mmproj and quant-matched
draft sidecars inherit their main quant's tier. Sidecar picking
mirrors the server's find_best_sibling ranking (deepest directory,
exact quant tag, closest bit depth).

Also port detectToolUseSupport (infers tool-calling support from a
chat template) and the browser get_info fallback helper.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 548acc2ffd ui : move catalog url to models-discover constants
Address review follow-up: the llama.app catalog endpoint belongs to
the models-discover feature, not the HF constants. Use Number() for
shard index parsing and name the UD-quant prefix segment lookup.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 7e63e9e346 ui : move HF constants to constants and enums modules
Address review on the HF data layer:

- replace the HfModelSort / SidecarForm / sibling entry type string
  unions with enums (HfModelSort, SidecarForm, HfEntryType)
- move URLs, query params, regexes, limits, retry settings, shard
  file conventions, tag tokens and formatting units into a dedicated
  huggingface.constants.ts; reuse the existing PATH_SEPARATOR
- drop the task label / pipeline icon / library display maps: the
  discover UI only presents GGUF models, so keep the task tags for
  logic use only (parseTags)
- drop the hardcoded curated model list; the discover dialog gets its
  default list from the llama.app /v1/catalog.json endpoint, which is
  an acceptable online-only source since the feature requires internet
  access anyway

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier ddb8d219e1 ui : follow MODEL_ID regex rename in huggingface service
Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier 4aa7063be3 ui : add huggingface hub data layer
Add HuggingFaceService for browsing and searching GGUF models on the
HF Hub: catalog/model search, model details, repo file tree, raw
README fetch, and the llama.app model catalog. Includes GGUF file
analysis helpers - extractQuantMeta (quant token plus sidecar type
and its form, prefix or suffix), shard collapsing, quant bit-depth
lookup, and download/size/likes formatting.

Add the HF API types and the curated model list shown in the
Discover Models sidebar.

Assisted-by: pi
2026-09-01 00:22:57 +02:00
Aleksander Grygier f6f90463be ui : lowercase reasoning capability and use sidecar enums in tests
Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 4c93cef4f5 ui : address review on model id parsing
Use lowercase values for the sidecar enums so the value doubles as
the filename token, derive the sidecar regexes from the enum values,
and rename the MODEL_ID regex keys to the _REGEX suffix used by the
rest of the constants files. Replace the tools capability magic
string with ModelCapability.TOOL_USE.

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier c1cbd5e277 ui : parse sidecar types in model ids
Add ModelDraftSidecar / ModelAuxSidecar enums with a ModelSidecar
union type; mmproj is the only auxiliary sidecar (single member,
covers vision and audio input). Add SIDECAR_PREFIX/SUFFIX_RE regex
matching the server's filename conventions, and type guards +
enum-file-token helpers in model-id.constants.ts.

Extend parseModelId to detect sidecar filename tokens (mtp-, mmproj-,
etc) and expose isDraftSidecar / isAuxSidecar / sidecarFromFileToken
helpers. Add ModelCapability.TOOL_USE with icon/label/flag mappings.

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 4b5449a445 ui : add format:files script for formatting changed files only
Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 9a8519788b ui : rename model API types to mirror endpoints
Drop the Router prefix from client-side API types; names now map
directly to the /models endpoint family (load/unload/download/list).
Merge ApiModelListResponse into ApiModelsListResponse (same endpoint
shape in both modes) and remove the duplicate ModelsService.listRouter().

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Aleksander Grygier 2853c1816c ui : remove dead ChatFormActionAddMcpSubmenu component
Unreferenced leftover from the pre-dialog MCP design; references
context props that no longer exist and breaks svelte-check.

Assisted-by: pi
2026-09-01 00:22:56 +02:00
Buğra Özgürsoy 458681e1d5 metal : add fa-vec tunings for M1 Ultra (#28088)
* metal : add fa-vec tunings for M1 Ultra

* metal : move M1 Ultra tunings after M1 Max section

* metal : remove duplicate blank line
b10729
2026-08-31 23:47:27 +02:00
ynankani e4b9af007b CUDA: XOR swizzle flash attn K,V smem fp16 tiles (#25635)
* CUDA: XOR swizzle flash attn  K,V smem fp16 tiles

Signed-off-by: ynankani <[email protected]>

* Fix use 64bit generic pointer instead of 32bit shared pointer

Signed-off-by: ynankani <[email protected]>

* fix shared memory race in FA on DGX Spark

* Handle corener case

Signed-off-by: ynankani <[email protected]>

* Add swizzle test cases and gate sync for swizzled path only

Signed-off-by: ynankani <[email protected]>

* gate CUDA PTX

Signed-off-by: ynankani <[email protected]>

* offset calculation specific for swizzle branch

Signed-off-by: ynankani <[email protected]>

* Reafctor code

Signed-off-by: ynankani <[email protected]>

* Refactor FA swizzle ldmatrix if/else into helpers (K row/col, V offset)

Signed-off-by: ynankani <[email protected]>

* rebase and update test case args

Signed-off-by: ynankani <[email protected]>

* Allow swizzle for non-pow2 shapes, for which nbatch_2%32==0

Signed-off-by: ynankani <[email protected]>

---------

Signed-off-by: ynankani <[email protected]>
b10728
2026-08-31 22:18:01 +02:00
Georgi Gerganov ab0b3bd3c8 metal : add concat support for quantized types (#28116)
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-0731
b10727
2026-08-31 23:16:04 +03:00
BartowskiandGeorgi Gerganov 85c55223ca AVX2: Speed up large batch size prompt processing of IQ models (#27402)
* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <[email protected]>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991e.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <[email protected]>
b10726
2026-08-31 14:33:50 -04:00
Georgi Gerganov 2a74817f93 metal : add top-k radix implementation (#28073)
Assisted-by: DeepSeek-v4-Flash-0731
2026-08-31 21:31:53 +03:00
itsnotoger 2d8d612e4c kv-cache : optimize restoring non-contiguous cells (#27991)
* kv cache : batch state restore scatter reads per contiguous run

When restoring state into non-contiguous destination cells (e.g. a
prompt-cache snapshot into a fragmented ring), state_read_data issued
one small copy per KV cell - ~1.4M copies of a few KiB each for a
40k+ token restore, taking 25-63 s on the CUDA backend.

The snapshot stores cell rows in cell order, so a maximal run of
consecutive destination indices maps to one contiguous block and can
be restored with a single copy. Precompute the runs once and use them
in all three scatter loops (K, V, transposed V). Byte-identical.

The on-device reader copies with a byte cursor when the read and
write chunking differs, so the batched reads are safe for it as well.
Batching makes equal tensor counts with a different split reachable
(save ranges [2,1] vs restore runs [1,2]); the next commit teaches the
reader's 1:1 path to fall back to the byte cursor in that case.

Verified in a production setup: 1,363,616 copies / 25-63 s -> 224
copies / 221-424 ms for the same restores (42,603 cells, 4 runs).

Assisted-by: Claude Code (unsloth/qwen3.8-27b)

* context : fall back to the byte cursor when read and write chunking differ

the on-device reader copies saved state back with a 1:1 copy by tensor
index whenever the write and read sides recorded the same number of
tensors, guarded by a per-tensor size assert.

equal tensor counts do not imply equal chunking: a state restore may
batch its reads per contiguous run of destination cells while the save
used per-range reads, so both sides can record two tensors that split
the same data differently, and the assert aborts in all builds.

compare the per-tensor sizes and only take the 1:1 path when the
chunking actually matches, otherwise fall through to the existing
byte-cursor copy. both sides enumerate the same logical data in the
same order, so the cursor copy is well-defined across tensor
boundaries.

Assisted-by: Claude Code (unsloth/qwen3.8-27b)

* tests : cover state restore scatter reads on host and on-device paths

decode the same prefix on two sequences, interleaving the seq 0 cells
between the seq 1 cells, so the seq 1 cells are isolated from each
other in the kv cache (three cells, two saved ranges). save the seq 1
state, free the interleaved seq 0 cells, and restore: the destination
is then non-contiguous (two runs), and the restore-side chunking has
the same tensor count as the save-side with a different split, so the
scatter path is batched per contiguous run and the on-device reader's
byte-cursor fallback is exercised.

the restored state is saved again on the host and compared byte for
byte with the first save: the blob is serialized in sequence cell
order, so the two saves are identical if and only if the scatter
restore wrote exactly the same KV content. this documents the
byte-identical guarantee of the run-batched scatter reads.

one test per io backend: the host (CPU) path and the on-device path.

Assisted-by: Claude Code (unsloth/qwen3.8-27b)
b10724
2026-08-31 19:49:58 +03:00
Hongqiang Wang 010be9683a opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (#26438)
* opencl: Q4_K/Q5_K mul_mv N_DST 4->8 on Intel for 2x activation reuse

* opencl: Q4_K mul_mm 8x8 tile fot Intel

* opencl: Q5_K mul_mm 8x8 tile for Intel

* opencl: Q4_K mul_mv N_DST 8->16 for Intel
b10723
2026-08-31 08:56:22 -07:00
Pascal 774ee0e200 ui: copy the displayed text of grouped agentic responses (#27832)
* ui: copy the displayed text of grouped agentic responses

Agentic sessions render as a single entry anchored on the first
assistant turn, whose content is typically just the first tool call,
so the copy button wrote an empty string to the clipboard. Derive the
text sections of the whole session and copy them joined, matching the
visible response. Plain messages keep the previous behavior.

* const
2026-08-31 17:48:43 +02:00
8e53fcefd2 webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation (#28045)
* webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation

* chore : improve code readability

Co-authored-by: Sigbjørn Skjæret <[email protected]>

---------

Co-authored-by: Stanisław Szymczyk <[email protected]>
Co-authored-by: Sigbjørn Skjæret <[email protected]>
b10721
2026-08-31 16:04:38 +02:00
Jaden_Mach f8dbcd6189 ROCm: add radix TOP_K for long rows (#27466)
* ROCm: add radix TOP_K for long rows
b10720
2026-08-31 15:00:04 +02:00
Niklas Wenzel 5d4a3be26d metal : add fa-vec tunings for M1 (#28078) b10719 2026-08-31 13:58:55 +02:00
ynankani 41ef91f7c8 CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token (#27621)
* CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were resticted to 1 token

Signed-off-by: ynankani <[email protected]>

* Address review comments

Signed-off-by: ynankani <[email protected]>

* Add SWIGLU_CLAMP case to multi-token moe fusion

Signed-off-by: ynankani <[email protected]>

---------

Signed-off-by: ynankani <[email protected]>
b10718
2026-08-31 19:22:28 +08:00
Neo Zhang a32af33de2 sycl : Enhance to get the free memory of Intel GPU (#27968)
* enhance get mem info by l0 an SYCL API

* remove debug code, format the code

* update SYCL.md for GGML_SYCL_GET_MEM_API
b10717
2026-08-31 13:33:02 +03:00
Sigbjørn Skjæret 580e88d8b7 ci : add check for unzip (#28082) 2026-08-31 12:17:51 +02:00
662a0b0121 spec : fuse the DFlash encoder into the KV cache injection (#27310)
* dflash : fuse the encoder into the KV injection decode

The encoder is a single fc + norm, but running it as a separate
llama_encode forced a device-to-host round trip of its output before the
injection decode could re-upload it, plus a second graph build per
round. Fold the encoder into the decoder's embd branch and feed the
target features directly to one llama_decode.

Assisted-by: Claude Fable

* nit

* Apply batched suggestions from code review

Co-authored-by: Ruixiang Wang <[email protected]>

* Fix missing references from renaming

---------

Co-authored-by: Sigbjørn Skjæret <[email protected]>
Co-authored-by: Ruixiang Wang <[email protected]>
b10715
2026-08-31 11:19:20 +02:00