Qwen3-VL-8B LoRA R16 SFT โ€” OmniFall All

Rank-16 SFT LoRA adapter for Qwen/Qwen3-VL-8B-Instruct, fine-tuned jointly on the OmniFall staged, synthetic, and in-the-wild training partitions. It predicts one of 16 human-activity labels from a video clip.

OmniFall

Input prompt

Place the video first in the user message, then append the following text. No system prompt is required.

Role:
You are an expert Human Activity Recognition (HAR) specialist.

Task:
Analyze the video clip and classify the primary action being performed.
Assign exactly one label from the allowed list below.

Note that the clip may contain more than one action. If this is the case,
focus on classifying the action in the first part of the clip, not the entire clip.
Example: Clips shows a person jumping and then falling. The correct label is jump.

Allowed Labels:
- walk
- fall
- fallen
- sit_down
- sitting
- lie_down
- lying
- stand_up
- standing
- other
- kneel_down
- kneeling
- squat_down
- squatting
- crawl
- jump

Output Format:
Respond with 'The best answer is: <class_label>' where <class_label> is one of the allowed labels.
Downloads last month
33
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for MoritzM00/qwen3-vl-8b-lora-r16-sft-omnifall-all

Adapter
(216)
this model

Paper for MoritzM00/qwen3-vl-8b-lora-r16-sft-omnifall-all