Tom Zicarelli, Author at reactive music

February 4, 2014September 5, 2014

ep-4yy13 DSP – week 2

new composition tools

from various artists

Tristan Jehan, Brian Wittman, and Paul Lamere: http://echonest.com/ and http://musicmachinery.com – infinite jukebox, remix.js, various others
Katja Vetter: http://www.katjaas.nl/slicejockey/slicejockey.html local: slicejockey2test2/slicejockey2test2.pd
Karlheinz Essl: RTClib local: RTC-lib_50-2/put content into patches/RTC-lib/Harmony/infinity-row.maxpat
Paul Nasca: extreme sound stretching http://musicmachinery.com/2013/11/26/scary-and-stretched/
Dinahmoe: Plink http://labs.dinahmoe.com/plink/
Celemony: Melodyne http://www.celemony.com/en/start – Minor version of “Bohemian Rhapsody”: http://www.youtube.com/watch?v=voca1OyQdKk
Brian Eno – Bloom
Naila Burney: Fictional dialog
Vocaloid
Mark Durham: formant synthesis http://sounddesignwithmax.blogspot.com/2013/05/i.html
Takahiko Tsuchiya: Probability based drum sequencer https://reactivemusic.net/?p=9233
Andreas Witsch: SoundEmotion2 https://reactivemusic.net/?p=9225
sorting sound
RJDJ
Twitter streaming API: https://reactivemusic.net/?p=5786
Echonest segment player
Ableton Live looper

tools that make tools

Alex Harker – impulse response tools.
Return of Pluggo: https://reactivemusic.net/?p=9636

February 4, 2014June 25, 2014

Notes: Chatbots in Conversation

update 6/2014 – Now part of the Internet sensors projects: https://reactivemusic.net/?p=5859

original post

They can talk with each other… sort of.

Last spring I made a project that lets you talk with chatbots using speech recognition and synthesis. https://reactivemusic.net/?p=4710.

Yesterday I managed to get two instances of this program, running on two computers, using two chatbots, to talk with each other, through the air. Technical issues remain (see below). But there were moments of real interaction.

In the original project, a human pressed button in Max to start and stop recording speech. This has been automated. The program detects and records speech, using audio level sensing. The auto-recording sensor turns on a switch when the level hits a threshold, and turns off after a period of silence. Threshold level and duration of silence can be adjusted by the user. There is also a feedback gate that shuts off auto-record while the computer is converting speech to text, and ‘speaking’ a reply.

technical issues

The Google speech API has difficulty with some of the voices used by the Mac OS speech synthesizer. We’ll need to experiment to find which voices produce accurate results.
The overall levels produced by the builtin Macbook speakers is not quite enough to achieve clear communication. The auto-recorder missed the onset of speech sometimes. One solution would be to insert a click to trigger the recorder, just before the speech synthesizer begins the actual speech. Or to use external speakers, or a secondary “wired” connection.
It would be nice to have menus of chatbots and voices. Also to automate the start of a new conversation thread.
The button to start the audio detector had to be operated by key-press because pushing the trackpad on a MacBook makes too much noise and always triggers the audio level detector.
Occasionally a chat bot would deliver a long response, or one containing a web address. These were problematic for recognition and synthesis.