ansaurus

Question

Verify image sequence

Answer 1

+1 A:

Perhaps you can find a way to get a binary copy of the image data of each frame in a variable. Hash that data (md5?) and store each of the hashes. Then you can see if you've ever seen that hash before. If you haven't, it's a new frame.

philfreo 2010-10-21 01:48:13

Answer 2

+2 A:

I think you have a few issues with this:

Not all image sequences [videos] are equal [but many are similar]
Where is your data coming from?
How will you repesent the data related to your viewings?
Size of the data

Issue #1:

Many images can differ slightly by compression, water marking, missing frames, and adding clips. I would suggest sampling the video. For example you may want to consider sub-sampling small sections of the images in the video. Additionally, to avoid noisy images and issues with lossely compression algorithms. You may want to consider grayscaling the frames sampled, and doing a gaussian blur. [Guassian because its "more natural" (short answer)] Once you have enough sub samples to where you have a good confidence of similarity to the video then store it in a database. With the samples you can hash them, or store them to do a % similarity later.

Issue #2

Your datasource is going to influence the tool kits, and libraries that you use. I would suggest keeping this simple [keep it with gifs and create a custom viewer, dont' try to write a browser plugin while developing your logic]

Issue #3

Using something like Postgres [if there are a lot of large sized objects] or SQLLite is highly suggested for indexing, storing, and recalling past meta data.

Issue #4

The size of the data will have a huge determination on recall, sampling, querying the database, etc.

Overall advice: Don't bite off more than you can handle at this stage. Start small and then grow.

Also take a look at Computer Vision algorithms for more help on the object representation/recall.

monksy 2010-10-23 04:25:00

Answer 3

+1 A:

The question itself is sure very interesting and challenging, however there are many practical issues as stated by @monksy.

The opportunist pragmatic in me would take a step back, look at the big picture and see if there is another way to solve the problem. For example, if you are building some kind of "image sharing community" and want to avoid duplicates in the database, you could do a simple md5 on the file (animated gifs on the web are usually always the same, it's rare that people modify them).

Another example: if you are analyzing scientific samples (like meteo sequences) it may be easier to directly embed some kind of hash in every file when generating them.

Francesco De Vittori 2010-10-23 10:22:07

Well you are correct... most of the time its just copied. However sometimes longer videos are cut into smaller ones.... my suggestion avoids this.

monksy 2010-10-23 20:39:07

Answer 4

+1 A:

This depends on wether you only want to know wether you've seen an absolutely identical movie again, or you also want to identify movies that are very similar but have been changed a bit (made lighter, have a watermark added, compression changed, etc.)

In the first case, just take any type of hash of the file and use that (because the file will be identical on the binary level.

In the second case (which I think is what you want) you have an interesting image processing problem on your hands. You could find yourself at the front-lines of image processing science with this if you'd want. If that is the case I suggest you start reading about SURF and OpenCV, and continue on from that.

If you want to match very similar, but not identical videos, and don't want to go the ultra-robus scientific route then I'd suggest the following process:

Do the gaussian blur you already do.
Divide each image into a few equally sized rectangles (you'd have to test for the best number, but I'd suggest you start with 9.
For each rectangle in each frame compute the full-colour histogram, then find the most occurring colour in that rectangle. This gives you 9*20 = 180 numbers. This is the "fingerprint" of this movie.
Find the most similar fingerprint in your database, if it is similar enough you already know about it, otherwise you don't.

Step 4 is a bit vague because I'm not really into this field. You are currently using an MD5 hash as a sort of fingerprint, but this is unsuitable in this case because slight differences in the input of a good cryptographic hashing function produce very large differences in the hash. This will mean that two very similar frames will have a totally different MD5 hash, so from the hash you'd never know they were similar.

As long as speed of database lookups is not an issue I'd just go for the sum of square differences as a measure of fingerprint similarity, and set a threshold on that to identify equal movies. However, this is not very fast for huge datasets, and in those cases you'd probably need to transform your fingerprint to something that will allow you to find similar fingerprints faster. One thing you could do here is start by selecting all known movies with very similar average colour for the entire video, then from that select the movies that have very similar average colour in each frame, and in the ones that remain at that point do the full rectangle-by-rectangle fingerprint match. But I'm sure there are even faster options for matching 180 numbers.

jilles de wit 2010-10-26 08:58:28

ansaurus

tags:

views:

answers:

Verify image sequence

Problem

Problem shaping

Limitations and uses

My effort

Gaussian blur

Data input, hashing, grayscaleing and lightboxing

Source example

related questions