如何看待AMD下一代Zen6架构主流桌面级 (MSDT) 处理器倒行逆施,疑似去掉集成显卡?

集成显卡其实很有用的。

跑个 AI 小任务不成问题。


用板载 iGPU 玩 Ai, 物尽其用并不丢人


D:\_LLAMA\llama.cpp>D:\_LLAMA\llama.cpp\build\bin\Release\llama-cli.exe -m g:\Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf   --list-devices
Available devices:
  CUDA0: NVIDIA GeForce RTX 5080 (16302 MiB, 14985 MiB free)
  Vulkan0: NVIDIA GeForce RTX 5080 (15977 MiB, 15209 MiB free)
  Vulkan1: AMD Radeon(TM) Graphics (99010 MiB, 94059 MiB free)
  BLAS: OpenBLAS (0 MiB, 0 MiB free)


D:\_LLAMA\llama.cpp>D:\_LLAMA\llama.cpp\build\bin\Release\llama-cli.exe -m g:\Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf   -dev Vulkan1










llama.cpp\build\bin\Release\llama-cli.exe -m g:\Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf --list-devices

Available devices:

CUDA0: NVIDIA GeForce RTX 5080 (16302 MiB, 14985 MiB free)

Vulkan0: NVIDIA GeForce RTX 5080 (15977 MiB, 15209 MiB free)

Vulkan1: AMD Radeon(TM) Graphics (99010 MiB, 94059 MiB free)

BLAS: OpenBLAS (0 MiB, 0 MiB free)


llama-cli.exe -m g:\Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf -dev Vulkan1


D:\_LLAMA\llama.cpp>D:\_LLAMA\llama.cpp\build\bin\Release\
llama-cli.exe -m g:\Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf   -dev Vulkan1

Loading model...


▄▄ ▄▄
██ ██
██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                    ██    ██
                                    ▀▀    ▀▀

build      : b9570-3ac3c20c9
model      : Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf
modalities : text

available commands:
  /exit or Ctrl+C     stop or exit
  /regen              regenerate the last response
  /clear              clear the chat history
  /read <file>        add a text file
  /glob <pattern>     add text files using globbing pattern


> tell a jioke

[Start thinking]
Thinking Process:

1.  **Analyze the Request:**
    *   Input: "tell a jioke"
    *   Interpretation: "jioke" is a common typo or slang for "joke" (specifically, "joke" with an 'i' instead of 'o', or possibly referring to the brand "Jio" in a Punjabi context, but overwhelmingly likely just a typo for "joke").
    *   Task: Tell a joke.
    *   Tone: Lighthearted, humorous, helpful.

2.  **Determine the Content:**
    *   Need a general joke that is suitable for a broad audience.
    *   Avoid any potential offensive content.
    *   Keep it concise.

3.  **Drafting Jokes:**
    *   *Option 1 (Pun):* Why did the scarecrow win an award? Because he was outstanding in his field.
    *   *Option 2 (Wordplay):* What do you call a fake noodle? An impasta.
    *   *Option 3 (Classic):* I told my wife she was drawing her eyebrows too high. She looked surprised.
    *   *Option 4 (Tech):* Why did the programmer go broke? Because he used `==` instead of `=`.

4.  **Selecting the Best Option:**
    *   Option 1 is classic and safe.
    *   Option 4 is relatable for many, but Option 1 is universally understood.
    *   Let's go with a short, classic one.

5.  **Refining the Output:**
    *   Acknowledge the typo playfully? (Maybe not necessary, but nice).
    *   Deliver the joke clearly.

6.  **Final Decision:**
    *   Go with the scarecrow joke or a simple one-liner. Let's pick the "Why did the math book look sad?" one. It's visual and funny.
    *   Actually, let's stick to the classic scarecrow one. It's a crowd-pleaser. Or maybe a wordplay one.
    *   Let's do this one:
    *   "What do you call a bear with no teeth? A gummy bear."
    *   Or the programmer one since I am an AI.
    *   Let's go with the "math book" one.
    *   Wait, the user typed "jioke". Maybe they are Punjabi speaking? "Jioke" sounds like "Joke" but with an 'i'. Or maybe "Jio" (the company) + "oke"? No, it's almost certainly "joke".
    *   I'll provide a clean joke.

    *   *Joke:* Why did the scarecrow win an award? Because he was outstanding in his field.

7.  **Final Polish:**
    *   Keep it friendly.

    *   *Joke:* Why did the math book look sad?
    *   *Answer:* Because it had too many problems.

    *   *Alternative:* A riddle? No, just a joke.

    *   Let's go with a pun about "jioke" just in case? No, keep it standard.

    *   Okay, here's a good one:
    *   "I asked my wife to make me a sandwich. She came back with a peanut butter and jelly sandwich."
    *   "What do you call a fake noodle? An impasta."

    Let's use the impasta one. It's short and punchy.

    Wait, "jioke" might be a reference to "Jio" (the telecom company). Is there a specific joke? Probably not. It's likely just a typo. I will treat it as "joke".

    *Joke:*
    Why did the scarecrow win an award?
    Because he was outstanding in his field.

    Let's try to be a bit more interactive.

    Okay, final choice:
    "Why did the bicycle fall over? Because it was two-tired!"

    Actually, the "outstanding" one is the most reliable.

    Let's write it.

    Wait, I should check if "jioke" is slang for something specific.
    Search query simulation: "jioke meaning".
    It's primarily a typo for "joke". Sometimes used in typing speed tests or autocorrect errors.
    Okay, I'll just tell a joke.

    *Joke:*
    What do you call a fish with no eyes?
    Fsh.

    Okay, let's go with that one-liner. It's clean.

    *Revised Plan:*
    1. Acknowledge the request.
    2. Tell the joke.

    *Draft:*
    Here's one for you!
    Why did the scarecrow win an award?
    Because he was outstanding in his field.

    *Wait, let's try to be more creative.*
    How about this one:
    "I invented a new word! Plagiarism!"
    No, too dark.

    "Why don't scientists trust atoms? Because they make up everything!"
    This is a classic.

    Let's go with the Atom joke. It's safe and clever.

    *Final check:* Is there any reason "jioke" means something else?
    In some contexts, "Jio" is a brand. But "jioke" isn't a standard slang.
    I'll assume "joke".

    Let's deliver the Atom joke.

    Wait, I'll provide two jokes just in case? No, one good one is better.
    Let's pick the "Math book" one.

    "Why did the math book look sad? Because it had too many problems."
    That works.

    Okay, let's output.cw

[End thinking]

Why did the math book look sad?

Because it had too many **problems**! 😄

[ Prompt: 9.7 t/s | Generation: 7.5 t/s ]






@echo off

set MODEL=G:\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf

set EXE=D:\_LLAMA\llama.cpp\build\bin\Release\llama-server.exe

"%EXE%" --model "%MODEL%" --device Vulkan1 --host 0.0.0.0 --port 8033 --ctx-size 8192 -b 512 -ub 512 -ctk f16 -ctv f16 --threads 4 --threads-batch 4 --parallel 1 --n-gpu-layers 99 --flash-attn on --no-host --reasoning off --reasoning-budget 0 --reasoning-format none --chat-template-kwargs "{\"enable_thinking\":false}" --jinja


pause


@echo off
set MODEL=G:\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf
set EXE=D:\_LLAMA\llama.cpp\build\bin\Release\llama-server.exe
"%EXE%" --model "%MODEL%"  --device Vulkan1 --host 0.0.0.0 --port 8033 --ctx-size 8192  -b 512 -ub 512  -ctk f16 -ctv f16 --threads 4 --threads-batch 4 --parallel 1 --n-gpu-layers 99 --flash-attn on --no-host --reasoning off --reasoning-budget 0 --reasoning-format none --chat-template-kwargs "{\"enable_thinking\":false}"   --jinja 

pause





D:\whispCPP>.\build\bin\Release\whisper-server.exe -m models\ggml-large-v2-q5_0.bin --host 0.0.0.0  --port 8080  -dev 2
whisper_init_from_file_with_params_no_state: loading model from 'models\ggml-large-v2-q5_0.bin'
whisper_init_with_params_no_state: use gpu    = 1
whisper_init_with_params_no_state: flash attn = 1
whisper_init_with_params_no_state: gpu_device = 2
whisper_init_with_params_no_state: dtw        = 0
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 16302 MiB):
  Device 0: NVIDIA GeForce RTX 5080, compute capability 12.0, VMM: yes, VRAM: 16302 MiB
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 5080 (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2
ggml_vulkan: 1 = AMD Radeon(TM) Graphics (AMD proprietary driver) | uma: 1 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 32768 | int dot: 1 | matrix cores: none
whisper_init_with_params_no_state: devices    = 5
whisper_init_with_params_no_state: backends   = 4
whisper_model_load: loading model
whisper_model_load: n_vocab       = 51865
whisper_model_load: n_audio_ctx   = 1500
whisper_model_load: n_audio_state = 1280
whisper_model_load: n_audio_head  = 20
whisper_model_load: n_audio_layer = 32
whisper_model_load: n_text_ctx    = 448
whisper_model_load: n_text_state  = 1280
whisper_model_load: n_text_head   = 20
whisper_model_load: n_text_layer  = 32
whisper_model_load: n_mels        = 80
whisper_model_load: ftype         = 8
whisper_model_load: qntvr         = 1
whisper_model_load: type          = 5 (large)
whisper_model_load: adding 1608 extra tokens
whisper_model_load: n_langs       = 99
whisper_model_load:      Vulkan1 total size =  1080.10 MB
whisper_model_load: model size    = 1080.10 MB
whisper_backend_init_gpu: device 0: CUDA0 (type: 1)
whisper_backend_init_gpu: found GPU device 0: CUDA0 (type: 1, cnt: 0)
whisper_backend_init_gpu: device 1: Vulkan0 (type: 1)
whisper_backend_init_gpu: found GPU device 1: Vulkan0 (type: 1, cnt: 1)
whisper_backend_init_gpu: device 2: Vulkan1 (type: 2)
whisper_backend_init_gpu: found GPU device 2: Vulkan1 (type: 2, cnt: 2)
whisper_backend_init_gpu: using Vulkan1 backend
whisper_backend_init: using BLAS backend
whisper_init_state: kv self size  =   83.89 MB
whisper_init_state: kv cross size =  251.66 MB
whisper_init_state: kv pad  size  =    7.86 MB
whisper_init_state: compute buffer (conv)   =   35.67 MB
whisper_init_state: compute buffer (encode) =   55.35 MB
whisper_init_state: compute buffer (cross)  =   47.67 MB
whisper_init_state: compute buffer (decode) =  100.04 MB

whisper server listening at http://0.0.0.0:8080


 


whisper-server.exe -m models\ggml-large-v2-q5_0.bin --host 0.0.0.0 --port 8080 -dev 2



iGPU 不是不能用, 只是有点慢。



..

编辑于 2026-06-18 · 著作权归作者所有