使用Colab+免费T4运行测试OneJev, Mac Mini M4使用 llama.cpp有问题

William 2026-10-06 13:10 1

看到OneJev OneJev:全面开源的多模态决策模型, 在huggingface跑了几下, 觉得还可以, 但是huggingface太小气了,只能试几次就提示付费了. 那想到Colab下载快又免费, 于是自己动手了



只能跑4B, 占用13G显存, 注意原模型使用bf16, T4不支持bf16, 要使用float16, 不然速度慢几十倍




跑一次要3秒左右, 简单的全对, 全开源+支持图片视频, 微调也容易, 作者 ^-^:




在MacMiniM4上, 使用llama.cpp速度非常慢, 13个问题, 花费41秒, 然后分析了代码后发现会拼接成接口数据如下:


{
"prompt": [
{
"prompt_string": "<|im_start|>system\nApply the question to the state. Choose exactly one of the listed options. Respond with only its uppercase letter, with no explanation or reasoning.<|im_end|>\n<|im_start|>user\n<state>\n{\n \"task\": \"Inspect the phone screen\",\n \"screen\": \"<__media_OhkEztkU4W7OzHaOdypO5FbPEsdrfG2e__>\"\n}\n</state>\n\nQuestion: The screen is the iPhone home screen showing a grid of app icons\n\nOptions:\nA. yes: the statement is true / the answer is yes\nB. no: the statement is false / the answer is no\n\nAnswer with one letter: A, B.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
"multimodal_data": [
"..."
]
},
{
"prompt_string": "<|im_start|>system\nApply the question to the state. Choose exactly one of the listed options. Respond with only its uppercase letter, with no explanation or reasoning.<|im_end|>\n<|im_start|>user\n<state>\n{\n \"task\": \"Inspect the phone screen\",\n \"screen\": \"<__media_OhkEztkU4W7OzHaOdypO5FbPEsdrfG2e__>\"\n}\n</state>\n\nQuestion: The Alipay (支付宝) app icon is visible on this screen\n\nOptions:\nA. yes: the statement is true / the answer is yes\nB. no: the statement is false / the answer is no\n\nAnswer with one letter: A, B.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
"multimodal_data": [
"..."
]
},
{
"prompt_string": "<|im_start|>system\nApply the question to the state. Choose exactly one of the listed options. Respond with only its uppercase letter, with no explanation or reasoning.<|im_end|>\n<|im_start|>user\n<state>\n{\n \"task\": \"Inspect the phone screen\",\n \"screen\": \"<__media_OhkEztkU4W7OzHaOdypO5FbPEsdrfG2e__>\"\n}\n</state>\n\nQuestion: The WeChat (微信) app icon is visible on this screen\n\nOptions:\nA. yes: the statement is true / the answer is yes\nB. no: the statement is false / the answer is no\n\nAnswer with one letter: A, B.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
"multimodal_data": [
"..."
]
}
],
"n_predict": 1,
"cache_prompt": true,
"temperature": 0,
"n_probs": 64
}

一个问题就会生成一条prompt, 一条prompt花费3秒多, 13个问题就会花40多秒

现在只是简单分析, 没有搞明白为什么一prompt花费这么多, 为什么同一个图片多个prompt没有共享缓存, 每次都重新分析, 两个问题导致非常慢

最新回复 (5)
  • Zane你发财 10-06 13:11
    1楼

    佬友是,自己电脑上跑不了,所以搞这个办法?

  • William 楼主 10-06 13:14
    2楼

    Colab简单下载快, 简单快速试一下

  • 今天不coding 10-06 13:23
    3楼

    可以,支持一下。正好可以学习一下这个

  • Astla 10-06 14:36
    4楼

    佬友太强了 支持一波 ^-^ ^-^colab很强大

  • William 楼主 10-07 01:43
    5楼

    更新了在Mac Mini m4上部署,对于llama.cpp不了解, 希望能帮分析一下

* 帖子来源Linux.do
返回