Donsanti92
Registered Member
- Jun 27, 2016
- 87
- 35
In the past few months, my main focus has been on YouTube's Content ID system. While I believe I may have found a solution, it is not yet solid-proof. However, those who are knowledgeable about AI may be able to help us with the final piece of the puzzle.
We know that YouTube scans content and generates fingerprints for both the audio and video. These fingerprints are then stored in their database to compare against future duplicates. I am aware that creators often work with DRM companies, which act as intermediaries between creators and YouTube's Content ID system. These companies upload the creators' content and check for duplicates. If they find any, they share the revenue collected.
I have read some research papers on fingerprinting methods, which mention that the science still cannot detect some modifications when they are properly modified. I decided to investigate this further.
In my case, I only have problems with audio. Therefore, I concentrated on building my own fingerprinting tool to scan my modified audio sample against the original audio sample. Here are my results:
In these experiments I had two videos, which I will call X and Y.
1st Experiment, I used Video X:
I compared the original and modified audio samples using a fingerprinting technique. The results showed that the modified audio was 20% similar to the original. However, when I uploaded the video to a test channel, I received a claim when the video reached 100k views.
2nd Experiment, I used Video Y:
I then tested video Y with another layer of modifications on the audio, reducing the similarity to only 3% compared to the original audio. However, once again, the DRM company claimed my video after it received more than 130k views.
3rd Experiment I used Video X:
In my third experiment, I used the same video & audio from 1st experiment but added an extra modification to it, BOOM the claim was not made!
But the weird thing is that the similarity percentage with the original audio was 25%. This puzzled me, as the video received more than 700k views, they should've claimed it already!
Quick summary:
They claimed me when I had 20% similarity.
They claimed me when I had 3% similarity.
They didn’t claim me when I had 25% similarity which is the highest I had!
I continued with these experiments with different videos, but to sum up, it is still a hit-and-miss situation (30% success rate).
If anyone can provide more insight, it would be greatly appreciated.
PS: I know some people may think this is a lot of work and may ask why I can't use 100% original content. I have other channels where I use 100% original content, but blackhat STILL ROCKS my friend.
We know that YouTube scans content and generates fingerprints for both the audio and video. These fingerprints are then stored in their database to compare against future duplicates. I am aware that creators often work with DRM companies, which act as intermediaries between creators and YouTube's Content ID system. These companies upload the creators' content and check for duplicates. If they find any, they share the revenue collected.
I have read some research papers on fingerprinting methods, which mention that the science still cannot detect some modifications when they are properly modified. I decided to investigate this further.
In my case, I only have problems with audio. Therefore, I concentrated on building my own fingerprinting tool to scan my modified audio sample against the original audio sample. Here are my results:
In these experiments I had two videos, which I will call X and Y.
1st Experiment, I used Video X:
I compared the original and modified audio samples using a fingerprinting technique. The results showed that the modified audio was 20% similar to the original. However, when I uploaded the video to a test channel, I received a claim when the video reached 100k views.
2nd Experiment, I used Video Y:
I then tested video Y with another layer of modifications on the audio, reducing the similarity to only 3% compared to the original audio. However, once again, the DRM company claimed my video after it received more than 130k views.
3rd Experiment I used Video X:
In my third experiment, I used the same video & audio from 1st experiment but added an extra modification to it, BOOM the claim was not made!
But the weird thing is that the similarity percentage with the original audio was 25%. This puzzled me, as the video received more than 700k views, they should've claimed it already!
Quick summary:
They claimed me when I had 20% similarity.
They claimed me when I had 3% similarity.
They didn’t claim me when I had 25% similarity which is the highest I had!
I continued with these experiments with different videos, but to sum up, it is still a hit-and-miss situation (30% success rate).
If anyone can provide more insight, it would be greatly appreciated.
PS: I know some people may think this is a lot of work and may ask why I can't use 100% original content. I have other channels where I use 100% original content, but blackhat STILL ROCKS my friend.