<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://jaekookang.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://jaekookang.github.io/" rel="alternate" type="text/html" /><updated>2025-01-23T06:18:48+00:00</updated><id>https://jaekookang.github.io/feed.xml</id><title type="html">Speech &amp;amp; Technology</title><subtitle>Speech &amp; Technology</subtitle><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><entry><title type="html">Articulatory Data Explorer</title><link href="https://jaekookang.github.io/posts/2025/01/speech-data/" rel="alternate" type="text/html" title="Articulatory Data Explorer" /><published>2025-01-23T00:00:00+00:00</published><updated>2025-01-23T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2025/01/articulatory-data-explorer</id><content type="html" xml:base="https://jaekookang.github.io/posts/2025/01/speech-data/"><![CDATA[<div style="text-align: center;">
    <img src="/images/speech-dataset_250123.png" />
</div>

<p><br /></p>

<p>I’m currently working on this web documentation (gitbook) for curating and organizing articulatory/speech dataset. Articulatory data is very rare, difficult to try out, but useful in research and make various applications in clinical and educational areas.</p>

<div>
  <p>Link: <a href="https://articulatory.gitbook.io/dataset" target="_blank">https://articulatory.gitbook.io/dataset</a></p>
</div>

<p>250123 jkang</p>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="articulatory data" /><category term="EMA" /><category term="rtMRI" /><category term="ultrasound" /><category term="speech data" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">AI Service for Digital Art in EdTech</title><link href="https://jaekookang.github.io/posts/2024/10/ai-digitalart/" rel="alternate" type="text/html" title="AI Service for Digital Art in EdTech" /><published>2024-10-01T00:00:00+00:00</published><updated>2024-10-01T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2024/10/ai-digitalart</id><content type="html" xml:base="https://jaekookang.github.io/posts/2024/10/ai-digitalart/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="/images/iscreamarts/iscreamart_ai_image1.png" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>The recent advance in AI and LLMs also had a tremendous impact on the education sector. For example, AI-based courseware and products are already available in the market focusing on helping children solve math problems, learning writing, lanugage learning, and so on. However, it has been somewhat still limited in terms of AI-based application in art, especially digital art education.</p>

<p>When I worked in i-Scream arts where making digital art education platform (“Art Bonbon”) and providing coaching online as a director of AI research, I had the opportunity to combine AI technologies with the digital art education. I made contribution to the development of three AI-based services for digital art education, which include:</p>

<h1 id="1-ai-based-progress-report-ai-report">1. AI-based Progress Report (“AI Report”)</h1>
<ul>
  <li>Aim: Generate an automatic analysis report for children’s art-learning progress to parents.</li>
  <li>Description: a weekly or monthly report based on children’s learning data (audio/drawing/metadata) using GPT.</li>
  <li>Outcome: Included in the digital art education platform as a service since 2023.
<img src="/images/iscreamarts/AI리포트_이미지1.jpg" alt="" /></li>
</ul>

<p><br /></p>

<h1 id="2-ai-based-drawing-diagnostic-tests">2. AI-based Drawing Diagnostic Tests</h1>
<ul>
  <li>Aim: Build a system for an automatic diagnostic drawing test for psychological assessment (digital art therapy)</li>
  <li>Description: Two drawing diagnostic tests were developed using custom AI models and MLLMs, measuring the amount of stress, coping strategy and emotionality in children.</li>
  <li>Outcome: Drawing tests have been developed and beta-tested successfully and will be launched at selected schools in Korea in 2024.</li>
  <li>Links:
    <ul>
      <li><a href="https://kidd.co.kr/news/234204">인터뷰: ‘비 오는 날’ 그리면 AI가 심리 분석해 줘</a></li>
      <li><a href="https://www.i-screammall.co.kr/goods/view?no=1158963">공식: 아이스크림몰 AI그림심리검사 구매</a></li>
      <li><a href="https://www.lecturernews.com/news/articleView.html?idxno=158293">기사: 디지털 기반 단체 그림검사 제공으로 아동미술 교육 혁신 이끈다</a>
<!-- ![](/images/iscreamarts/AI그림심리검사-이미지1.png)
![](/images/iscreamarts/AI그림심리검사-이미지2.png)
![](/images/iscreamarts/AI그림심리검사-이미지3.png) --></li>
    </ul>
  </li>
</ul>

<p><img src="/images/iscreamarts/psych-pitr-1.jpg" alt="" /></p>
<div style="display: flex; text-align: center; max-width: 30%; height: auto;">
  <img src="/images/iscreamarts/psych-image1.png" /> 
  <img src="/images/iscreamarts/psych-image2.png" /> 
  <img src="/images/iscreamarts/psych-image3.png" /> 
</div>

<p><br /></p>

<h1 id="3-ai-based-sketch-tool-ai-sketch">3. AI-based Sketch Tool (“AI Sketch”)</h1>
<ul>
  <li>Aim: Help children in the initial process of drawing a sketch using AI</li>
  <li>Description: ‘AI Sketch’ function was included on Art Bonbon’s platform (drawing tool) such that when children take or upload a photo, AI model extracts sketches for them.</li>
  <li>Outcome: Two main functions were developed (photo-based AI Sketch &amp; generative image-based AI Sketch using Stable Diffusion)</li>
</ul>

<div style="display: inline; max-width: 100%; height: auto;">
  <div style="display: flex; text-align: center; max-width: 40%">
    <img src="/images/iscreamarts/sketch-image1.jpeg" /> 
    <img src="/images/iscreamarts/sketch-image2.jpeg" /> 
    <img src="/images/iscreamarts/sketch-image3.jpeg" /> 
  </div>
</div>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="EdTech" /><category term="Digital Art" /><category term="AI-based service" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Articulatory Data Processor</title><link href="https://jaekookang.github.io/posts/2021/08/artic-data-processor/" rel="alternate" type="text/html" title="Articulatory Data Processor" /><published>2021-01-08T00:00:00+00:00</published><updated>2021-01-08T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2021/08/artic-data-extractor%20copy</id><content type="html" xml:base="https://jaekookang.github.io/posts/2021/08/artic-data-processor/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="/images/iscreamarts/모음이미지1.png" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>This repository includes a procedure for processing data features either from articulation and acoustics. This step precedes the procedure described <a href="https://github.com/jaekookang/Articulatory-Data-Extractor">here</a>. This is part of unpublished work of Jaekoo Kang’s dissertation. Note that data files used here are not uploaded here. You may need to request the use of data if necessary, but it is not guaranteed.</p>

<h1 id="overview">Overview:</h1>
<ul>
  <li>Outlier removal</li>
  <li>Excluding short tokens (&lt;step_size for formant tracking)</li>
  <li>Visualizations for articulatory and acoustic data (data_plots)</li>
  <li>Saving the result file (.csv) under data_processed</li>
</ul>

<h1 id="note">Note:</h1>
<ul>
  <li>
    <p>Data files will not be provided in this repo. You can follow steps described in <a href="https://github.com/jaekookang/Articulatory-Data-Extractor">link1</a> and <a href="https://github.com/jaekookang/Python-EMA-Viewer">link2</a> to generate data files.</p>
  </li>
  <li>
    <p>GitHub: <a target="_blank" href="https://github.com/jaekookang/Articulatory-Data-Processor">https://github.com/jaekookang/Articulatory-Data-Processor</a></p>
  </li>
</ul>

<div style="display: flex; max-width: 100%; height: auto;">
  <div style="text-align: center; justify-content: center;">
    <img style="margin: 10px;" src="/images/iscreamarts/모음이미지2.png" />
    <img style="margin: 10px;" src="/images/iscreamarts/모음이미지3.png" />
    <img style="margin: 10px;" src="/images/iscreamarts/모음이미지4.png" />
  </div>
</div>

<p><br /></p>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="speech production" /><category term="electromagnetic articulography" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Articulatory Data Extractor</title><link href="https://jaekookang.github.io/posts/2021/08/artic-data-extractor/" rel="alternate" type="text/html" title="Articulatory Data Extractor" /><published>2021-01-08T00:00:00+00:00</published><updated>2021-01-08T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2021/08/artic-data-extractor</id><content type="html" xml:base="https://jaekookang.github.io/posts/2021/08/artic-data-extractor/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="https://raw.githubusercontent.com/jaekookang/Articulatory-Data-Extractor/master/png/top.png" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>This repository includes a procedure for extracting articulatory features with the corresponding acoustic features from the electromagnetic articulography (EMA) dataset. Currenetly, this procedure is only optimized for the Haskins IEEE EMA dataset (Link). Support for different datasets will be considered if necessary.</p>

<h1 id="features">Features</h1>
<ul>
  <li>Articulatory features: pellet sensor coordinates (horizontal, vertical)</li>
  <li>Acoustic features: formant frequencies (F1, F2, F3) and f0</li>
  <li>meta: speaker, utterance, word, phone etc.</li>
</ul>

<h1 id="data-input-output">Data input-output</h1>
<ul>
  <li>Input: EMA data (*.mat or *.pkl)</li>
  <li>
    <p>Output: articulatory and acoustic data features (*.pkl or *.csv)</p>
  </li>
  <li>GitHub: <a target="_blank" href="https://github.com/jaekookang/Articulatory-Data-Extractor?tab=readme-ov-file">https://github.com/jaekookang/Articulatory-Data-Extractor?tab=readme-ov-file</a></li>
</ul>

<p><br /></p>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="speech production" /><category term="electromagnetic articulography" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Implementation of Invertible Neural Networks (INNs) based on Ardizzone et al. (2019) using TensorFlow2 + Keras</title><link href="https://jaekookang.github.io/posts/2020/11/inn-basic/" rel="alternate" type="text/html" title="Implementation of Invertible Neural Networks (INNs) based on Ardizzone et al. (2019) using TensorFlow2 + Keras" /><published>2020-12-14T00:00:00+00:00</published><updated>2020-12-14T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2020/11/inn-basic</id><content type="html" xml:base="https://jaekookang.github.io/posts/2020/11/inn-basic/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="https://github.com/jaekookang/invertible_neural_networks/blob/master/result/gauss_mixture.gif?raw=true" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>This repository includes the code example of the invertible neural networks (Ardizzone et al., 2019) implemented using TensorFlow2 with Keras. Two complementary coupling layers were implemented and toy examples were provided similar to the paper. The current code is largely based on the original PyTorch implementation by the authors, but simplified for the easier understanding than the original code and for the personal use.</p>

<ul>
  <li>GitHub: <a target="_blank" href="https://github.com/jaekookang/invertible_neural_networks">https://github.com/jaekookang/invertible_neural_networks</a></li>
  <li>Colab: <a target="_blank" href="https://colab.research.google.com/github/jaekookang/invertible_neural_networks/blob/master/colab_example_gaussian_mixture.ipynb"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab" /></a></li>
</ul>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="speech production" /><category term="invertible neural networks" /><category term="speech motor control" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Estimating “Good” variability in Speech Production using Invertible Neural Networks</title><link href="https://jaekookang.github.io/posts/2012/08/inn-issp2020/" rel="alternate" type="text/html" title="Estimating “Good” variability in Speech Production using Invertible Neural Networks" /><published>2020-12-14T00:00:00+00:00</published><updated>2020-12-14T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2012/08/inn-issp2020</id><content type="html" xml:base="https://jaekookang.github.io/posts/2012/08/inn-issp2020/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="/images/iscreamarts/inn-image1.png" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>Variability is inherent in skilled human motor movements. Playing a piano or riding a bicycle requires skilled coordination of motor elements, such as arms and legs, to achieve a motor goal. Although the movements are skillful, the positions of the motor elements are not exactly the same regardless of how many times they are repeated or executed (“repetition without repetition”; Bernstein, 1967). This variability in the form of repeated limb movements can be understood as an informative biological feature in the human motor system due to its underlying structure and regularity (Latash et al., 2002; Riley &amp; Turvey, 2002; Sternad, 2018; Whalen &amp; Chen, 2019), which previously had been disregarded as noise. One such structure of the skilled motor movements is that it is highly synergistic and flexibly organized when decomposed into “good” and “bad” parts of variability (i.e., the uncontrolled manifold approach or the UCM; Latash et al., 2002; Scholz &amp; Schöner, 1999, 2014). Whether variability in speech production can also be decomposed into the same principle, however, has been rarely examined to date. Specifically, this project aims to focus on the “good” part of variability in speech production and explore the use of invertible neural networks as a quantitative approach to understand how “good” variability is structured and can be learned by these neural-net models.</p>

<p><br /></p>

<p><a target="_blank" href="https://jaekookang.me/issp2020/">International Seminar on Speech Production (ISSP) Project Website</a></p>
<p>Below is iframe of the webpage.</p>

<p><br /></p>

<div style="display: flex; width: 100%; height: auto;">
<iframe src="https://jaekookang.me/issp2020/" height="1000px" frameborder="0" scrolling="yes" allowfullscreen="" style="width: -webkit-fill-available;">
</iframe>
</div>

<!-- width="500" height="300"  -->]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="speech production" /><category term="invertible neural networks" /><category term="speech motor control" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Linear Transformations on the Web</title><link href="https://jaekookang.github.io/posts/2020/05/linear-transformations/" rel="alternate" type="text/html" title="Linear Transformations on the Web" /><published>2020-05-01T00:00:00+00:00</published><updated>2020-05-01T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2020/05/linear-transformations</id><content type="html" xml:base="https://jaekookang.github.io/posts/2020/05/linear-transformations/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="https://github.com/jaekookang/linear_transformations/blob/master/public/preview.gif?raw=true" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>This is an interactive webapp for visualizing basic linear transformations. Based on Dr. Lauren K Williams’ amazing JavaScript Applets, this adds a bit more of tweak and modifications for the personal use and experimentations.</p>

<p>More functionalities (e.g., random shape and cascading transformations) were added in addition to Dr. Lauren’s original web app
Code were bundled using Vite</p>

<ul>
  <li>Demo: <a target="_blank" href="https://linear-transformations.netlify.app">https://linear-transformations.netlify.app <a href="https://app.netlify.com/sites/linear-transformations/deploys"><img src="https://api.netlify.com/api/v1/badges/7bdf6862-d28a-4618-a34e-cf27b386043f/deploy-status" alt="Netlify Status" /></a></a></li>
  <li>GitHub: <a target="_blank" href="https://github.com/jaekookang/linear_transformations?tab=readme-ov-file">https://github.com/jaekookang/linear_transformations?tab=readme-ov-file</a></li>
</ul>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="speech production" /><category term="invertible neural networks" /><category term="speech motor control" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Forward-mapping of speech articulators</title><link href="https://jaekookang.github.io/posts/2018/11/speech-vis/" rel="alternate" type="text/html" title="Forward-mapping of speech articulators" /><published>2020-03-30T00:00:00+00:00</published><updated>2020-03-30T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2018/11/artic-forward-map</id><content type="html" xml:base="https://jaekookang.github.io/posts/2018/11/speech-vis/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="/images/articulation/ucm_example2_1.gif" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p><br /></p>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="automatic speech recognition" /><category term="speech visualization" /><category term="realtime visualization" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Articulatory Data Visualization - EMA/MVIEW</title><link href="https://jaekookang.github.io/posts/2018/11/artic-vis-ema/" rel="alternate" type="text/html" title="Articulatory Data Visualization - EMA/MVIEW" /><published>2020-02-19T00:00:00+00:00</published><updated>2020-02-19T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2018/11/artic-vis-ema</id><content type="html" xml:base="https://jaekookang.github.io/posts/2018/11/artic-vis-ema/"><![CDATA[<h1 id="biteblock-experiment-preliminary">Biteblock Experiment (preliminary)</h1>
<p><img src="/images/articulation/biteblock1.png" alt="biteblock1" />
<br /></p>

<p><img src="/images/articulation/biteblock2.png" alt="biteblock1" />
<br /></p>

<p><img src="/images/articulation/biteblock3.png" alt="biteblock1" />
<br /></p>

<h1 id="had-normal">“had” (normal)</h1>
<p><img src="/images/articulation/had.gif" alt="biteblock1" /></p>

<h1 id="had-with-biteblock">“had” (with biteblock)</h1>
<p><img src="/images/articulation/had_biteblock.gif" alt="biteblock1" /></p>

<h1 id="had-shadowing">“had” (shadowing)</h1>
<p><img src="/images/articulation/had_shadow.gif" alt="biteblock1" /></p>

<h1 id="had-shadowing-with-biteblock">“had” (shadowing with biteblock)</h1>
<p><img src="/images/articulation/had_shadow_biteblock.gif" alt="biteblock1" /></p>

<p><br /></p>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="electromagnatic articulography" /><category term="speech visualization" /><category term="speech motor control" /><summary type="html"><![CDATA[Biteblock Experiment (preliminary)]]></summary></entry><entry><title type="html">Speech Inversion</title><link href="https://jaekookang.github.io/posts/2020/01/speech-inversion/" rel="alternate" type="text/html" title="Speech Inversion" /><published>2020-01-26T00:00:00+00:00</published><updated>2020-01-26T00:00:00+00:00</updated><id>https://jaekookang.github.io/posts/2020/01/speech-inversion</id><content type="html" xml:base="https://jaekookang.github.io/posts/2020/01/speech-inversion/"><![CDATA[<div style="display: inline; max-width: 100%; height: auto;">
  <div style="text-align: center;">
    <img src="/images/articulation/speech-inversion.png" />
  </div>
  <!-- <p style="text-align: center; font-size: smaller">
    <a target="_blank" href="https://www.hankyung.com/article/2022111596825">파블로아트컴퍼니, 미국 실리콘밸리서 아트봉봉 소개 | 한국경제</a>
  </p> -->
</div>

<p><br /></p>

<p>The task of speech inverion from acoustics to articulation is tested and visualized.</p>

<ul>
  <li>Paper: <a target="_blank" href="https://www.researchgate.net/publication/309755936_Development_of_articulatory_estimation_model_using_deep_neural_network">You, H., Yang, H., Kang, J., Cho, Y., Hwang, S. H., Hong, Y., Cho, Y., Kim, S., &amp; Nam, H. (2016). Development of articulatory estimation model using deep neural network. Phonetics and Speech Sciences, 8(3), 31–38. http://dx.doi.org/10.13064/KSSS.2016.8.3.031
</a></li>
</ul>

<p><br /></p>

<h1 id="test">Test</h1>
<p><img src="/images/articulation/spinv_lr_2020-01-26.gif" alt="" /></p>

<p><br /></p>]]></content><author><name>Jaekoo Kang</name><email>jaekoo.jk@gmail.com</email></author><category term="speech inversion" /><category term="lstm" /><summary type="html"><![CDATA[]]></summary></entry></feed>