9654 x 2 + 630内存 + 4060ti 跑GLM-5.3-Flash-Q4_K_M

dtfour09 2026-08-30 23:43 1

3.35.792.741 D srv operator (): all results received, terminating stream

3.35.792.757 D srv operator (): http: streamed chunk: data: [DONE]


3.35.792.797 D srv operator (): http: stream ended


[ Prompt: 7.1 t/s | Generation: 2.5 t/s ]


Exiting...

3.35.797.233 D que start_loop: processing new tasks

3.35.797.235 D que process_new_: terminate

3.35.797.552 I srv operator (): operator (): cleaning up before exit...

3.35.814.072 I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |

3.35.814.081 I common_memory_breakdown_print: | - CUDA0 (RTX 4060 Ti) | 8187 = 0 + ( 7436 = 5316 + 0 + 2120) + 751 |

3.35.814.082 I common_memory_breakdown_print: | - Host | 175712 = 174922 + 673 + 116 |

3.35.814.239 D ~llama_context: CUDA0 compute buffer size is 2120.0234 MiB, matches expectation of 2120.0234 MiB

3.35.814.241 D ~llama_context: CUDA_Host compute buffer size is 116.4122 MiB, matches expectation of 116.4122 MiB


^-^


个人还是别玩了。


这是Windows 跑的,Linux可以会快一点点。

最新回复 (11)
  • 快乐的出帆 08-30 23:45
    1

    这个模型,体积多大?

  • dtfour09 楼主 08-30 23:47
    2

    @快乐的出帆 #1 大概占用190g内存。应该是180b的.


    320b的量化版本。

  • NicholasRobert 08-30 23:52
    3

    这体量个人实在玩不起 ^-^

  • GLM 08-30 23:55
    4

    这么贵的cpu,也只能跑 2.5 t/s ?

  • dtfour09 楼主 08-30 23:58
    5

    @GLM #4 显卡太拉了,另外是双路的拉跨点。


    最好是大显存 + gpu。

  • GLM 08-31 00:00
    6

    @dtfour09 #5 要不你纯cpu跑看看?感觉纯CPU跑比RTX 4060 Ti还快

  • dtfour09 楼主 08-31 00:05
    7

    @GLM #6 纯cpu更慢。 ^-^

  • chenhaha 08-31 00:14
    8

    这速度瓶颈还是显卡

  • dtfour09 楼主 08-31 00:18
    9

    @chenhaha #8 起码激活参数要够18b。

  • 快乐的出帆 08-31 00:20
    10

    512GB的M5U的Mac适合

  • dtfour09 楼主 08-31 00:23
    11

    @快乐的出帆 #10 卡的问题,显存不够,激活参数太差了。

    +1个 其实 200g内存 + 一个5090会快很多。

* 帖子来源NodeSeek
返回