TRECVID 2020 : Call for Participation in the Video Retrieval Evaluation Benchmark [Feb. 2020 – Nov. 2020]

CALL FOR PARTICIPATION in the 2020 TREC VIDEO RETRIEVAL EVALUATION (TRECVID 2020)
February 2020 - November 2020
Conducted by the National Institute of Standards and Technology (NIST) with additional funding from other US government agencies.
I n t r o d u c t i o n:
The TREC Video Retrieval Evaluation series (trecvid.nist.gov) promotes progress in content-based analysis of and retrieval from digital video via open, metrics-based evaluation. TRECVID is a laboratory-style evaluation that attempts to model real world situations or significant component tasks involved in such situations. In its 20th annual evaluation cycle TRECVID will evaluate participating systems on 6 different video analysis and retrieval tasks using various types of real world datasets. Below is the main datasets to be used in 2020 across the 6 proposed tasks.
D a t a:
In TRECVID 2020 NIST will use at least the following data sets:
      * Vimeo Creative Commons Collection (V3C)             The V3C is a large-scale video dataset that has been collected from high-quality        web videos with a time span over several years in order to represent true videos        in the wild. It consists of 28,450 videos with a duration of 3,801 hours in total.        The first part of this dataset (V3C1) has been used by the Video Browser Showdown        (VBS) 2019 and the Ad-Hoc Video Search (AVS) task at TRECVID 2019 as well.        For both campaigns V3C1 will serve as a basis over three years (VBS 2019-2021        and TRECVID 2019-2021). V3C1 contains 1,000 hours of video content and approximately one million        shots that were created by the authors of the dataset using the open-source multimedia        retrieval engine Cineast. A subset of approx. 2000 clips from the second part (V3C2) of the V3C       collection will be used as testing data for the Video-to-Text (VTT) task in 2020.              * IACC.3
      The IACC.3 was introduced in 2016 and consists of approximately 4600 Internet        Archive videos (144 GB, 600 h) with Creative Commons licenses in MPEG-4/H.264       format with duration ranging from 6.5 min to 9.5 min and a mean duration       of almost 7.8 min. Most videos will have some metadata provided by the       donor available e.g., title, keywords, and description.
      * BBC EastEnders
      Approximately 244 video files (totally 300GB, 464 hours) with       associated metadata, each containing a week's worth of BBC EastEnders       programs in MPEG-4/H.264 format.
      * Twitter Vine videos           Approximately 8,000 6 sec video clips URLs from the public Twitter stream of Vine videos        have been human annotated by video captions from 2016-2019. These Vine videos will be provided       as development data for participants of the Video-to-Text (VTT) task.
       * Gatwick and i-LIDS MCT airport surveillance video
      The data consist of about 150 hours obtained from airport       surveillance video data (courtesy of the UK Home Office). The       Linguistic Data Consortium has provided event annotations for       the entire corpus. The corpus was divided into development and       evaluation subsets. Annotations for 2008 development and test       sets are available.
      * VIRAT dataset          The VIRAT Video Dataset is a large-scale surveillance video dataset designed to        assess the performance of activity detection algorithms in realistic scenes.        The dataset was collected to facilitate both detection of activities and to localize        the corresponding spatio-temporal location of objects associated with activities from        a large continuous video. The VIRAT dataset are closely aligned with real-world video        surveillance analytics.
       * LADI dataset
      The Low Altitude Disaster Imagery (LADI) dataset is hosted as part of the AWS Public Dataset program        and will be available to participants of the DSDI task as development data. It consists of over 20,000+        annotated images, each at least 4 MB in size. The annotated features were selected based on a recommendation        from the public safety community. In total there are 31 features across 5 categories. The dataset was        collected between 2015 - 2019 during major natural disaster events (e.g. hurricanes, floodings, earthquakes)        across several USA states. The lower altitude criteria is intended to further distinguish the LADI dataset        from satellite or "top down" datasets and to support development of computer vision capabilities        for small drones operating at low altitudes. A minimum image size was selected to maximize the efficiency        of the crowd source workers. For more information about LADI, please refer to the        github organization.
 T a s k s:
In TRECVID 2020 NIST will evaluate systems on the following tasks using the [data] indicated:
     * AVS: Ad-hoc Video Search (automatic, manually-assisted, relevance feedback) [V3C1]
      The Ad-hoc search task started in TRECVID 2016 and will continue in 2020        to model the end user search use-case, who is looking for        segments of video containing persons, objects, activities, locations, etc.,        and combinations of the former. Given about 30 multimedia topics created at        NIST, return for each topic all the shots which meet the video need expressed        by it, ranked in order of confidence. Although all evaluated submissions will be       for automatic runs, Interactive systems will have the opportunity to       participate in the Video Browser Showdown (VBS) in 2021 using        the same testing data (V3C1).
     * ActEV: Activities in Extended Video [VIRAT]              ActEV is a series of evaluations to accelerate development of robust, multi-camera,        automatic activity detection algorithms for forensic and real-time alerting applications.        ActEV is an extension of the annual TRECVID Surveillance Event Detection (SED) evaluation        where systems will also detect, and track objects involved in the activities. Each evaluation        will challenge systems with new data, system requirements, and/or new activities.
           * INS: Instance search (interactive, automatic) [BBB EastEnders] 
      An important need in many situations involving video collections       (archive video search/reuse, personal video organization/search,       surveillance, law enforcement, protection of brand/logo use) is to       find more video segments of a certain specific person, object,       or place, given a visual example. A new query type started in 2019        asking systems to retrieve specific persons doing specific actions.        A set of defined actions with various image/video examples will be given        and each topic will include few examples (image and video)        of a person and ask systems to find that person doing one of the defined actions.
    * VTT: Video to Text Description [Vimeo Creative Commons Collections (V3C2)]
      Automatic annotation of videos using natural language text descriptions        has been a long-standing goal of computer vision. The task involves        understanding of many concepts such as objects, actions, scenes, person-object        relations, temporal order of events and many others. In recent years there have        been major advances in computer vision techniques which enabled researchers to        start practically to work on solving such problem.        Given a set of short video clips and number of reference sets of text descriptions,        systems are asked to work and submit results for two subtasks.The core "Description Generation" subtask requires       systems to automatically generate a text description (1 sentence) for each video clip.       An optional "Matching and Ranking" subtask requires systems to       return for each video a ranked list of the most likely text       description that correspond (was annotated) to the video from each of the reference sets.
    * VSUM: Video Summarization [BBC Eastenders Soap Opera]       An important need in many situations involving video collections (archive video search/reuse,        personal video organization/search, movies, tv shows, etc.) is to summarize the video in order        to reduce the size and concentrate the amount of high value information in the video track.        In 2020 we begin a new video summarization track in TRECVID in which the task is to       summarize the major life events of specific characters over a number of weeks of programming        on the BBC Eastenders TV series. Typically, three characters will be chosen for this task every year,        and summaries of their major life events must be between the selected period of the show, which will        be specified to participants in advance of the task.
    * DSDI: Disaster Scene Description and Indexing [Low Altitude Disaster Imagery (LADI)]
      Computer vision capabilities have rapidly been advancing and are expected to become an important        component to incident and disaster response. However, the majority of computer vision capabilities        are not meeting public safety needs, such as support for search and rescue, due to the lack        of appropriate training data and requirements. In response, the organizers developed a dataset of images       collected by the Civil Air Patrol of various natural disasters. Two key distinctions are the low altitude and        oblique perspective of the imagery and disaster-related features, which are rarely featured in computer vision        benchmarks and datasets. This task invites researchers to work on this new domain to develop new capabilities       and close the gap in performance to essentially label short video clips with the correct disaster-related feature(s).
              In addition to the data, TRECVID will provide uniform scoring procedures, and a forum for organizations interested in comparing their approaches and results.
Participants will be encouraged to share resources and intermediate system outputs to lower entry barriers and enable analysis of various components' contributions and interactions.
 *************************************************** * You are invited to participate in TRECVID 2020 * ***************************************************
The evaluation is defined by the Guidelines. A draft version is available: http://www-nlpir.nist.gov/projects/tv2020/index.html and further feedback input from the participants are welcomed till April,2020.
You should read the guidelines carefully before applying to participate in one or more tasks:  Guidelines
 P l e a s e   n o t e:   1) Dissemination of TRECVID work and results other than in the (publicly available) conference proceedings is welcomed, but the conditions of participation specifically preclude any advertising claims based on TRECVID results.
2) All system output and results submitted to NIST are published in the Proceedings or on the public portions of TRECVID web site archive.
3) The workshop is open only to participating groups that submit results for at least one task and to selected government personnel from sponsoring agencies and data donors.
4) Each participating group is required to submit before the November workshop a notebook paper describing their experiments and results. This is true even for groups who may not be able to attend the workshop.
5) It is the responsibility of each team contact to make sure that information distributed via the call for participation and the tv20.list@list.nist.gov email list is disseminated to all team members with a need to know. This includes information about deadlines and restrictions on use of data.
6) By applying to participate you indicate your acceptance of the above conditions and obligations.
 There is a tentative schedule for the tasks included in the Guidelines webpage: Schedule
 W o r k s h o p   f o r m a t
Plans are for a 2 and half days workshop at NIST in Gaithersburg, Maryland - just outside Washington, DC. Confirmation and details will be provided to participants as soon as available.
The TRECVID workshop is used as a forum both for presentation of results (including failure analyses and system comparisons), and for more lengthy system presentations describing retrieval techniques used, experiments run using the data, and other issues of interest to researchers in information retrieval and computer vision. As there is  a limited amount of time for these presentations, the evaluation coordinators and NIST will determine which groups are asked to speak and which groups will present in a poster session. Groups that are interested in having a speaking slot during the workshop will be asked to submit a short abstract before the workshop describing the experiments they performed. Speakers will be selected based on these abstracts.
 H o w   t o   r e s p o n d   t o   t h i s   c a l l
Organizations wishing to participate in TRECVID 2020 must respond to this call for participation by submitting an on-line application by 1 April (the earlier the better).  Only ONE APPLICATION PER TEAM please, regardless of how many organizations the team comprises.
*PLEASE* only apply if you are able and fully intend to complete the work for at least one task. Taking the data but not submitting any runs threatens the continued operation of the workshop and the availability of data for the entire community.
Here is the application URL:   http://ir.nist.gov/tv-submit.open/application.html
You will receive an immediate automatic response when your application is received. NIST will respond with more detail to all applications submitted before the end of March.  At that point you'll be given the active participant's userid and password, be subscribed to the tv20.list email discussion list, and can participate in finalizing the guidelines as well as sign up to get the data, which is controlled by separate passwords.
 T R E C V I D   2 0 2 0   e m a i l   d i s c u s s i o n   l i s t
The tv20.list email discussion list (tv20.list@list.nist.gov) will serve as the main forum for discussion and for dissemination information about TRECVID 2020.  It is each participant's responsibility to monitor the tv20.list postings.  It accepts postings only from the email addresses used to subscribe to it.  At the bottom of the guidelines there is a link to an archive of past postings available using the active participant's userid/password.
 Q u e s t i o n s
Any administrative questions about conference participation, application format/content, subscriptions to the tv20.list, etc. should be sent to george.awad at nist.gov.
 Best regards,
TRECVID 2020 organizers team 

Special Issue on “Facing emerging challenges in multimedia forensics”

;text-indent:0px;text-transform:none;white-space:normal;word-spacing:0px;word-wrap:break-word”>———————————-
Giulia Boato, PhD
Associate Professor 
DISI – University of Trento
tel: +39 0461 283193

pdf icon JINSCallforPapers.pdf
Read More »

FoNeS-IoT 2020 – EAI International Conference on Forthcoming Networks and Sustainability in the IoT Era

January 30th, 2020 Daniela Lopez de Luise
Web version

October 1 – 2, 2020 | Antalya, Turkey
Submission Deadline: April 10, 2020

SCOPE

The advancements in technology, application areas for advanced communication systems and development of new services, facilitate a tremendous growth of new devices and smart things that need to be connected to the Internet through a variety of wireless technologies. Parallel to this, new capabilities such as pervasive sensing, multimedia sensing, machine learning, deep learning, unmanned aerial vehicles, cloud and edge computing, energy efficiency/harvesting and computing power opens the way to new domains, services and business models beyond the traditional mobile Internet. The new areas in turn come with various requirements in terms of reliability, quality of service, and energy efficiency. These are only some examples of the challenges that are of interest to researchers in Forthcoming Networks and Sustainability in the IoT Era (FoNeS-IoT).

We are pleased to invite you to submit your paper to FoNeS-IoT 2020. Submissions should be in English, following the Springer formatting guidelines. (see Initial Submission). We also invite Poster proposals and young PhD candidates to participate in Doctoral Consortium.
Submit Paper
Publications

All accepted and presented papers will be submitted for publishing via Springer to be made available through SpringerLink Digital Library. Proceedings are also included in EUDL (EAI open access digital library).

Proceedings will be submitted for inclusion in leading indexing services, including Ei Compendex, ISI Web of Science, Scopus, CrossRef, Google Scholar, DBLP, as well as EAI’s own EU Digital Library (EUDL).

FoNeS-IoT proceedings will be submitted for publication in:

Additionally, selected papers will be considered to be included in one of EAI Transactions and the Springer’s Mobile Networks and Applications (MONET) Journal (IF: 2.390).

eai community benefits
Granting you visibility and
fair review through
Community Review.
Credits counting towards your   EAI Index , membership ranks and global recognition.
Receive invaluable real-time feedback on your presentation on-site via EAI Compass.
Important dates

Full Paper Submission Deadline: April 10, 2020

Notification Deadline: May 22, 2020

Camera-ready deadline: July 30, 2020

Conference dates: October 1 – 2, 2020
Follow us!

facebook    You Tube   twitter 

ONION: peOple in laNguage, visIOn and the miNd

January 28th, 2020 Daniela Lopez de Luise

Second Call for Papers

ONION: peOple in laNguage, visIOn and the miNd

Workshop to be held at the 12th Edition of the Language Resources and Evaluation Conference, Palais du Pharo, Marseilles, France, on Saturday, May 16 2020.

https://onion2020.github.io/

We invite paper submissions for the first workshop on People in Language, Vision, and the Mind, which discusses how people, their bodies and faces as well as mental states are described in text. We are interested in contributions from diverse areas including language generation, language analysis, cognitive computing, affective computing.

Detailed Workshop goals

The workshop will provide a forum to present and discuss current research focusing on multimodal resources as well as computational and cognitive models aiming to describe people in terms of their bodies and faces, including their affective state as it is reflected physically. Such models might either generate textual descriptions of people, generate images corresponding to people’s descriptions, or in general exploit multimodal representations for different purposes and applications.  Knowledge of the way human bodies and faces are perceived, understood and described by humans is key to the creation of such resources and models, therefore the workshop also invites contributions where the human body and face are studied from a cognitive, neurocognitive or multimodal communication perspective. 

Human body postures and faces are being studied by researchers from different research communities, including those working with vision and language modeling, natural language generation, cognitive science, cognitive psychology, multimodal communication and embodied conversational agents. The workshop aims to reach out to all these communities to explore the many different aspects of research on the human body and face, including the resources that such research needs,  and to foster cross-disciplinary synergy.

The ability to adequately model and describe people in terms of their body and face is interesting for a variety of language technology applications, e.g., conversational agents and interactive multimodal narrative generation, as well as forensic applications in which people need to be identified or their images generated from textual or spoken descriptions. 

Such systems need resources and models where images associated with human bodies and faces are coupled with linguistic descriptions, therefore the research needed to develop them is placed at the interface between vision and language research. 

At the same time, this line of research raises important ethical questions, both from the perspective of data collection methodology and from the perspective of bias detection and avoidance in models trained to process and interpret human attributes.

By focussing on the modelling and processing of physical characteristics of people, and the ethical implications of this research, the workshop will explore and further develop a particular area within visual and language research. Furthermore, it will foster novel cross-disciplinary knowledge by soliciting contributions from different fields of research. By attempting to bring results from the cognitive and neurocognitive fields to the attention of the HLT community, it is also in line with the “Language and the Brain” hot topic of LREC 2020.

Relevant topics

We are inviting short and long papers reporting original research, surveys, position papers, and demos. Authors are strongly encouraged to identify and discuss ethical issues arising from their work, insofar as it involves the use of image data or descriptions of people.

Relevant topics include, but are not limited to, the following ones:
  • Datasets of facial images, as well as body postures, gestures and their descriptions
  • Methods for the creation and annotation of multimodal resources dedicated to the description of people
  • Methods for the validation of  multimodal resources for descriptions of people
  • Experimental studies of facial expression understanding by humans
  • Models or algorithms for automatic facial description generation
  • Emotion recognition by humans
  • Multimodal automatic emotion recognition from images and text
  • Subjectivity in face perception
  • Communicative, relational and intentional aspects of head pose and eye-gaze
  • Collection and annotation methods for facial descriptions
  • Coding schemes for the annotation of body posture and facial expression
  • Understanding and description of the human face and body in different contexts, including commercial applications, art, forensics, etc. 
  • Modelling of the human body, face and facial expressions for embodied conversational agents
  • Generation of full-body images and/or facial images from textual descriptions
  • Ethical and data protection issues related to the collection and/or automatic description of images of real people
  • Any form of bias in models which seek to make sense of human physical attributes in language and vision.

Important dates

Paper submission deadline: February 14, 2020
Notification of acceptance: March 13, 2020 
Camera ready Papers: April 2, 2020

Workshop: May 16, 2020 (afternoon)  

Submission guidelines

Short paper submissions may consist of up to 4 pages of content, while long papers may have up to 8 pages of content. References do not count towards these page limits.

All submissions must follow the LREC 2020 style files, which are available for LaTeX (preferred) and MS Word and can be retrieved from the following address: https://lrec2020.lrec-conf.org/en/submission2020/authors-kit/

Papers must be submitted digitally, in PDF, and uploaded through the online submission system here:

https://www.softconf.com/lrec2020/ONION2020/

The authors of accepted papers will be required to submit a camera-ready version to be included in the final proceedings. Authors of accepted papers will be notified after the notification of acceptance with further details.

Identify, Describe and Share your LRs!
Describing your LRs in the LRE Map is now a normal practice in the submission procedure of LREC (introduced in 2010 and adopted by other conferences). To continue the efforts initiated at LREC 2014 about “Sharing LRs” (data, tools, web-services, etc.), authors will have the possibility,  when submitting a paper, to upload LRs in a special LREC repository. This effort of sharing LRs, linked to the LRE Map for their description, may become a new “regular” feature for conferences in our field, thus contributing to creating a common repository where everyone can deposit and share data.
As scientific work requires accurate citations of referenced work so as to allow the community to understand the whole context and also replicate the experiments conducted by other researchers, LREC 2020 endorses the need to uniquely Identify LRs through the use of the International Standard Language Resource Number (ISLRN, www.islrn.org), a Persistent Unique Identifier to be assigned to each Language Resource. The assignment of ISLRNs to LRs cited in LREC papers  will be offered at submission time.
Organisers

Patrizia Paggio, University of Copenhagen and University of Malta, paggio@hum.ku.dk
Albert Gatt, University of Malta, albert.gatt@um.edu.mt
Roman Klinger, University of Stuttgart, roman.klinger@ims.uni-stuttgart.de

Programme committee 

Adrian Muscat, University of Malta
Andreas Hotho, University of Würzburg
Andrew Hendrickson, University of Tilburg
Catherine Pelachaud, Institute for Intelligent Systems and Robotics, UPMC and CNRS 
Costanza Navarretta, CST, University of Copenhagen
David Hogg, University of Leeds
Diego Frassinelli, University of Stuttgart
Isabella Poggi, Roma Tre University
Jonas Beskow, KTH Speech, Music and Hearing
Jordi Gonzalez, Universitat Autònoma de Barcelona
Kristiina.Jokinen, National Institute of Advanced Industrial Science and Technology (AIST)
Mihael Arcan,  National University of Ireland, Galway
Raffaella Bernardi, CiMEC Trento
Sebastian Padó, University of Stuttgart


Patrizia Paggio

Professor
University of Malta
Institute of Linguistics and Language Technology
patrizia.paggio@um.edu.mt

Senior Researcher
University of Copenhagen
Centre for Language Technology
paggio@hum.ku.dk

ANNPR 2020 (9th IAPR TC3 Workshop on artificial NN’s in Pattern Recognition), Sept. 2nd-4th, 2020 in Winterthur, Switzerland

We are excited to announce ANNPR 2020, the 9th IAPR TC3 Workshop on Artificial Neural Networks in Pattern Recognition, which will be held from September 2nd-4th, 2020, at Zurich University of Applied Sciences ZHAW in Winterthur, Switzerland.

 

The workshop will act as a major forum for international researchers and practitioners working in all areas of neural network- and machine learning-based pattern recognition to present and discuss the latest research, results, and ideas in these areas. ANNPR is the biannual workshop organized by the Technical Committe 3 (IAPR TC3) on Neural Networks & Computational Intelligence of the International Association for Pattern Recognition (IAPR).

 

Among the previous editions of the workshop were ANNPR 2018 (Siena, Italy), ANNPR 2016 (Ulm, Germany), ANNPR 2014 (Montreal, Canada), ANNPR 2012 (Trento, Italy), ANNPR 2010 (Cairo, Egypt), ANNPR 2008 (Paris, France) and ANNPR 2006 (Ulm, Germany).

 

Program

 

The workshop will consist of keynote talks, several sessions for presentations of accepted papers, and a poster session. Keynote speakers are Jürgen Schmidhuber (tentative, The Swiss AI Lab IDSIA, Lugano, Switzerland), Naftali Tishby (Hebrew University of Jerusalem, Israel) and Bernd Feisleben (University of Marburg, Germany).

 

In addition, there will be a dedicated industry session featuring a sponsored keynote, applied research presentations and industry exhibits / booths / demos, which will provide networking opportunities. The social program consists of a welcome reception, an excursion and a conference dinner.

 

Topics

 

ANNPR 2020 invites papers that present original work in the areas of neural networks and machine learning oriented to pattern recognition, focusing on their algorithmic, theoretical, and applied aspects. Topics of interest include, but are not limited to:

 

Methodological Issues:

– Supervised, semi-supervised, unsupervised and reinforcement learning

– Deep learning & deep reinforcement learning

– Feed-forward, recurrent, and convolutional neural networks

– Generative models

– Interpretability & explainability of neural networks

– Robustness & generalization of neural networks

– Meta-learning, Auto-ML

 

Applications to Pattern Recognition:

– Image classification and segmentation

– Object detection

– Document analysis, e.g. handwriting recognition

– Sensor-fusion and multi-modal processing

– Biometrics, including speech and speaker recognition and segmentation

– Data, text, and web mining

– Bioinformatics and medical applications

– Industrial applications, e.g. quality control and predictive maintenance

 

Paper Submission (Deadline: May 1st, 2020)

 

Papers are invited to be submitted via EasyChair. There will be a peer review process before acceptance. The page limit is 12 pages. For instructions, see https://annpr2020.ch/cfp/.

 

Accepted papers will be published as a special volume of the Springer Lecture Notes in Artificial Intelligence (LNAI) series. For previous editions, see https://link.springer.com/conference/annpr.

 

Organization

 

The workshop is organized by the Institute of Applied Information Technology (InIT) at the ZHAW School of Engineering.

The program committee is listed at https://annpr2020.ch/organisation/

 

Chairs: Dr. Frank-Peter Schilling (chair), Prof. Dr. Thilo Stadelmann (co-chair)

 

Venue

 

ZHAW is one of the leading universities of applied sciences in Switzerland, with around 13 000 students and host to one of Europe’s first and largest dedicated research centers for Data Science. ZHAW’s School of Engineering is one of the leading Engineering Faculties in Switzerland. Our 13 institutes and centres guarantee superior-quality education, research and development with an emphasis on the areas of energy, mobility, information and health.

 

Winterthur is the sixth-largest city of Switzerland with around 110 000 inhabitants. It is located in eastern Switzerland, approximately 20 km from Zurich. The Zurich region is the hub of the digital economy in Switzerland, hosting companies like Google, NVIDIA, IBM and Facebook, which have a high natural affinity to neural networks and pattern recognition. It is also host to world famous universities such as ETH Zurich and University of Zurich UZH, as well as to several universities of applied sciences including the ZHAW. The Swiss AI Institute IDSIA in Lugano, as well as the Swiss Data Science Center SDSC by ETH Zurich and EPFL Lausanne complete the Artificial Intelligence and Machine Learning landscape in the region.

 

 

Best regards,

 

Dr. Frank-Peter Schilling (chair)

Prof. Dr. Thilo Stadelmann (co-chair)

 

Design by 2b Consult