Visual routine

A visual routine is a means of extracting information from a visual scene.

Shimon Ullman, in his studies on human visual cognition, proposed that the human visual system's task of perceiving shape properties and spatial relations is split into two successive stages: an early "bottom-up" state during which base representations are generated from the visual input, and a later "top-down" stage during which high-level primitives dubbed "visual routines" extract the desired information from the base representations.[1] In humans, the base representations generated during the bottom-up stage correspond to retinotopic maps (more than 15 of which exist in the cortex) for properties like color, edge orientation, speed of motion, and direction of motion. These base representations rely on fixed operations performed uniformly over the entire field of visual input, and do not make use of object-specific knowledge, task-specific knowledge, or other higher-level information.[2]

The visual routines proposed by Ullman are high-level primitives which parse the structure of a scene, extracting spatial information from the base representations. These visual routines are composed of a sequence of elementary visual operators specific to the task at hand. Visual routines differ from the fixed operations of the base representations in that they are not applied uniformly over the entire visual field --- rather, they are only applied to objects or areas specified by the routines.[1]

Ullman lists the following as examples of visual operators: shifting the processing focus, indexing a salient item for further processing, spreading activation over an area delimited by boundaries, tracing boundaries, and marking a location or object for future reference. When combined into visual routines, these elementary operators can be used to perform relatively sophisticated spatial tasks such as counting the number of objects satisfying a certain property, or recognizing a complex shape.[1]

قام عدد من الباحثين بتطبيق إجراءات بصرية لمعالجة صور الكاميرا، لأداء مهام مثل تحديد الجسم الذي يشير إليه الإنسان في صورة الكاميرا. [ 3 ] [ 4 ] [ 5 ] كما طبق الباحثون نهج الإجراءات البصرية على تمثيلات الخرائط الاصطناعية، لتشغيل ألعاب الفيديو ثنائية الأبعاد في الوقت الفعلي . في هذه الحالات، تم توفير خريطة لعبة الفيديو مباشرةً، مما أغنى عن الحاجة إلى التعامل مع مهام الإدراك في العالم الحقيقي مثل التعرف على الأشياء وتعويض الحجب .

مراجع

  1. 1 2 3 "الروتينات البصرية لأولمان، ورسومات تيكوتسو" (ملف PDF) .
  2. هوانغ، ج.؛ ويكسلر، هـ. (أبريل 2000). "روتينات بصرية لتحديد موقع العين باستخدام التعلم والتطور". معاملات IEEE في الحوسبة التطورية . 4 (1): 73-82 . doi : 10.1109/4235.843496 . ISSN 1089-778X . 
  3. جونسون، عضو البرلمان (أغسطس 1996). "إنشاء إجراءات بصرية تلقائيًا باستخدام البرمجة الجينية". وقائع المؤتمر الدولي الثالث عشر للتعرف على الأنماط . المجلد 1. الصفحات 951-956. doi : 10.1109/ICPR.1996.546164 . ISBN   978-0-8186-7282-8. S2CID 1701864 . 
  4. أستي، ماركو؛ روسي، ماسيمو؛ كاتوني، رولدانو؛ كابريل، برونو (1998-06-01). "إجراءات بصرية للمراقبة الآنية لسلوك المركبة". رؤية الآلة وتطبيقاتها . 11 (1): 16-23 . CiteSeerX 10.1.1.48.5736 . doi : 10.1007/s001380050086 . ISSN 0932-8092 . S2CID 25480778 .   
  5. راو، ساتياجيت. "الروتينات البصرية والانتباه" (ملف PDF) . مختبر علوم الحاسوب والذكاء الاصطناعي في معهد ماساتشوستس للتكنولوجيا .