WEBVTT

00:16.959 --> 00:19.459
Uh So for the team's awareness we met

00:19.469 --> 00:21.580
uh uh UB a little over a year ago , I

00:21.580 --> 00:23.747
guess , uh when he was doing a postdoc

00:23.747 --> 00:25.969
with uh Jan Lao following his uh phd at

00:25.969 --> 00:28.299
uh at Berkeley with uh uh Bruno

00:28.309 --> 00:31.030
Olshausen . Uh So we've got , we've got

00:31.040 --> 00:33.040
some nice alignment in our , in our

00:33.040 --> 00:34.707
interest , had good technical

00:34.707 --> 00:36.818
discussion , philosophical discussion

00:36.818 --> 00:38.929
at the time . Uh His , his focus on ,

00:38.929 --> 00:40.596
on representation learning in

00:40.596 --> 00:42.818
particular , the uh the intersection of

00:42.818 --> 00:44.596
uh self supervised learning and

00:44.596 --> 00:46.484
computational neuroscience really

00:46.484 --> 00:46.009
aligns . Well , I think with gel with

00:46.020 --> 00:48.242
folks in the , in the call here today .

00:48.242 --> 00:50.187
Anyway , I'll get out of the way .

00:50.187 --> 00:52.298
You'll be welcome . I appreciate it .

00:52.298 --> 00:54.353
Uh uh I appreciate again , you , you

00:54.353 --> 00:53.840
calling in and talking with us . Um I

00:53.849 --> 00:56.071
know this is uh uh exciting topic for ,

00:56.071 --> 00:58.889
for everybody . All right . Thanks for

00:58.900 --> 01:02.389
having me . And uh uh do you

01:02.400 --> 01:06.239
see this life ? Mhm . All

01:06.250 --> 01:08.361
right . So , uh it's a great pleasure

01:08.361 --> 01:11.010
to be here . And uh I'm Yu Chen and I

01:11.019 --> 01:14.919
just started a new research lab at uh

01:14.930 --> 01:18.290
UC Davis , the EC Department . But also

01:18.300 --> 01:20.522
I'm uh in the Neuroscience and Computer

01:20.522 --> 01:22.578
Science Department by Courtesy . And

01:22.578 --> 01:24.970
today I'm going to talk about um how

01:24.980 --> 01:27.769
can we um uh seek some principles of n

01:27.790 --> 01:31.080
supervised representation learning ? Um

01:32.089 --> 01:33.978
My lab is very new , right ? So I

01:33.978 --> 01:35.830
started uh uh just uh this uh uh

01:35.839 --> 01:39.610
January uh formally , so to establish

01:39.620 --> 01:42.059
that . So , so it's really a very new

01:42.069 --> 01:44.750
group . So today's talk would be about

01:44.819 --> 01:47.430
answer wise representation learning and

01:47.440 --> 01:49.607
uh on the top left corner , right . So

01:49.607 --> 01:51.662
there that , that's some of the most

01:51.662 --> 01:54.239
important works related in general . My

01:54.250 --> 01:56.860
work tend to consider uh representation ,

01:56.870 --> 01:58.981
learning from many different angles ,

01:58.981 --> 02:00.703
generating models is entangled

02:00.703 --> 02:03.029
representation , uh interpretation , uh

02:03.040 --> 02:06.010
representation for Motors efficiency ,

02:06.160 --> 02:09.289
uh geometric uh deep learning and real

02:09.300 --> 02:12.429
world uh A I systems and robustness .

02:12.440 --> 02:15.289
Um So , before I started with that ,

02:15.300 --> 02:17.649
maybe uh a little bit more about myself ,

02:17.660 --> 02:20.440
I have been trained by Ya Quin and

02:20.449 --> 02:22.869
Bruno , right ? So two very

02:22.880 --> 02:25.929
complementary um objectives . One is

02:25.940 --> 02:28.119
from deep learning engineering and one

02:28.130 --> 02:29.789
is computer uh computational

02:29.800 --> 02:31.911
neuroscience , right . So science and

02:31.911 --> 02:34.279
engineering fri so uh also a comp

02:34.539 --> 02:36.650
communication neuroscientist and also

02:36.650 --> 02:38.706
collab with uh collaborate with many

02:38.706 --> 02:42.520
different PIS Iman and Y JB Kiper

02:43.169 --> 02:46.429
and uh many of my fellow phd students

02:46.440 --> 02:48.710
and now many of them start to become uh

02:48.720 --> 02:52.139
PIS at different universities . Um um

02:52.149 --> 02:54.300
and some awesome gra uh graduate

02:54.309 --> 02:56.850
students and undergraduate students .

02:57.190 --> 03:00.330
Um My work used to be supported by uh A

03:00.339 --> 03:04.229
FRL um while I was uh at uh IIU and

03:04.240 --> 03:06.462
thanks for uh for your guests support .

03:06.462 --> 03:09.250
And uh before I joined um the UC Davis ,

03:09.259 --> 03:11.529
I spent a little bit of time at uh be R

03:11.539 --> 03:13.990
and ecs and computation uh and

03:14.000 --> 03:16.000
Neuroscience and NYU Data for CI uh

03:16.000 --> 03:18.559
data for uh Center for Data Science .

03:18.740 --> 03:21.020
Um And uh Redwood Theoretical

03:21.029 --> 03:24.679
Neuroscience . All right . So , so the

03:24.690 --> 03:28.070
current machine learning paradigm , we ,

03:28.080 --> 03:30.600
well , I prepared this slides uh last

03:30.610 --> 03:33.479
year and uh you know , things change

03:33.490 --> 03:35.712
very quickly , right ? So the current ,

03:35.712 --> 03:37.934
you know , when I say current is it was

03:37.934 --> 03:40.440
really about last year , last year's um

03:40.449 --> 03:43.039
major machine learning paradigm was

03:43.050 --> 03:45.320
super at learning . But this year I I

03:45.330 --> 03:48.113
suspect a lot of well , it becomes and

03:48.373 --> 03:50.595
learning becomes really important now .

03:50.595 --> 03:52.882
Um So , but that , that paradigm is

03:52.893 --> 03:54.893
basically you are given the image a

03:54.893 --> 03:57.473
sample and uh given uh uh labels ,

03:57.483 --> 03:59.650
right , supervised learning . And then

03:59.650 --> 04:01.650
we aim to learn a neural network or

04:01.650 --> 04:04.242
function that can maps our sample um to

04:04.253 --> 04:07.533
the target , right ? So um it works

04:07.542 --> 04:10.093
really so well and um started in 2

04:10.123 --> 04:13.945
2012 and uh Alex net , right ?

04:13.955 --> 04:16.196
So that's uh that , that was a really a

04:16.205 --> 04:18.876
breakthrough showing that basically a

04:18.885 --> 04:22.246
larger data and more um uh more data

04:22.295 --> 04:25.045
and more labels and larger network you

04:25.055 --> 04:27.235
can really uh solve , start to solve a

04:27.246 --> 04:29.545
lot of tasks better . But there are

04:29.555 --> 04:31.805
many issues um in , in this such a

04:31.816 --> 04:34.585
paradigm first , it requires a large

04:34.596 --> 04:37.600
numbers of LA labeled samples and

04:37.609 --> 04:41.010
that's a very manual labor intensive .

04:41.119 --> 04:44.000
Uh Second is that it would be very

04:44.010 --> 04:46.179
specialized to introduce human bias .

04:46.190 --> 04:49.100
For example , in this case , I , I um

04:49.109 --> 04:52.000
you guys can see my cursor , right ? OK ,

04:52.010 --> 04:55.130
cool . So in this case , I we label

04:55.140 --> 04:57.196
that as a dog but it really , it's a

04:57.196 --> 04:59.940
dog uh standing on the grassland . Um

04:59.950 --> 05:02.117
And there , there's some sunshine from

05:02.117 --> 05:04.600
the uh from a particular angle , right ?

05:04.609 --> 05:06.776
So , and when we label that as a dog ,

05:06.970 --> 05:09.192
we introduce human knowledge , but also

05:09.192 --> 05:13.089
human bias . And um uh further um

05:13.100 --> 05:15.720
sometimes this could be uh dangerous or

05:15.730 --> 05:18.290
unpractical for if we want to train a

05:18.299 --> 05:21.399
um uh um autonomous driving car , right ?

05:21.410 --> 05:23.619
So we need to crash that car so many

05:23.630 --> 05:25.963
times in order to tell the machine , OK ,

05:25.963 --> 05:28.529
that's a bad idea to not crash , right ?

05:28.540 --> 05:32.019
So um and then um it's a black box , we

05:32.029 --> 05:33.807
still don't understand what has

05:33.807 --> 05:36.190
happened in this case , right ? So how

05:36.200 --> 05:38.422
this neuronal accomplish that , right .

05:38.422 --> 05:40.910
So it's um it's a paradigm shift ,

05:40.920 --> 05:43.709
right ? So in the past , when we write

05:43.720 --> 05:46.390
a program , we we code it , right ? So

05:46.399 --> 05:48.399
we write the logic . Now we use the

05:48.410 --> 05:50.769
data to define and learn that logic ,

05:50.839 --> 05:52.895
right ? So that's a paradigm shift ,

05:52.895 --> 05:54.950
right ? So we use the data to define

05:54.950 --> 05:56.950
our program , but really , it's the

05:56.950 --> 05:58.783
black box , we still don't fully

05:58.783 --> 06:00.895
understand what's what has happened .

06:00.895 --> 06:03.117
And further it could be computationally

06:03.117 --> 06:05.283
intensive for both of the training and

06:05.283 --> 06:09.040
deployment . Uh compared to that ,

06:09.049 --> 06:10.882
if we are thinking about natural

06:10.882 --> 06:14.019
intelligence , um um a lot of the times

06:14.029 --> 06:16.980
we find that biological intelligence um

06:16.989 --> 06:19.220
uh demonstrate the capability far

06:19.230 --> 06:22.019
beyond any kind uh uh state of the art

06:22.029 --> 06:24.500
machine learning systems . And the

06:24.510 --> 06:27.440
difference is that um for biological

06:27.450 --> 06:29.519
intelligence , a lot of times how we

06:29.529 --> 06:31.820
learn is by uh looking around and

06:31.829 --> 06:35.260
interacting with the environment . Um

06:35.269 --> 06:37.519
so that we have some of the basic

06:37.529 --> 06:39.820
understanding of our environment . And

06:39.829 --> 06:42.570
for any particular task we can build on

06:42.579 --> 06:45.600
top of that um foundation , right ? So

06:45.609 --> 06:48.149
there are , there are two parts , one

06:48.160 --> 06:50.160
is the sensory , one is the motor ,

06:50.160 --> 06:52.160
right ? So sensory is the input and

06:52.160 --> 06:54.327
motor is the output . And we have this

06:54.327 --> 06:56.160
sensory motor loop . And today's

06:56.160 --> 06:58.399
majority of my talk would be on the

06:58.410 --> 07:00.709
sensory part . At the end , very end ,

07:00.720 --> 07:02.887
I will come back to the motor part and

07:02.887 --> 07:06.869
start to uh close the loop . Um So ,

07:06.880 --> 07:09.429
yeah , so and so was the representation

07:09.440 --> 07:12.010
learning really aims to mimic this uh

07:12.019 --> 07:14.420
uh particular paradigm and build

07:14.429 --> 07:16.651
machine learning systems that can learn

07:16.651 --> 07:19.329
automatically from the data and build

07:19.339 --> 07:21.630
transforms that can bring the data

07:21.720 --> 07:24.130
structure more explicit with the

07:24.140 --> 07:26.549
representation transform . In other

07:26.559 --> 07:28.760
words , we do not have that uh human

07:28.769 --> 07:31.329
label anymore . All we have is a lot of

07:31.339 --> 07:34.429
the samples and there are structures in

07:34.440 --> 07:36.496
those samples . And we want to build

07:36.496 --> 07:39.290
the um a machine learning algorithm

07:39.299 --> 07:41.739
that can discover that patterns from ,

07:41.750 --> 07:44.559
from , from the input samples a huge

07:44.570 --> 07:46.626
amount of them , right . So it's far

07:46.626 --> 07:48.848
from Gaussian noise , that's why we can

07:48.848 --> 07:51.014
learn . And the question is how can we

07:51.014 --> 07:53.500
actually find those structures um and

07:53.510 --> 07:56.170
build that transform ? And uh and in

07:56.179 --> 07:59.079
the Z is really our latent space and

07:59.089 --> 08:01.619
where the data structure is more

08:01.630 --> 08:05.600
explicit . So what is an answer

08:05.660 --> 08:08.670
towards the representation ? Um um uh

08:08.679 --> 08:11.160
in general , I would say that we want

08:11.170 --> 08:13.510
to build a transform uh that can uh

08:13.519 --> 08:15.519
transform our data into a new space

08:15.519 --> 08:17.670
such that the structure is more

08:17.679 --> 08:19.600
explicit , right ? So there are

08:19.609 --> 08:21.887
structure in the original data , right ?

08:21.887 --> 08:24.269
But it's just a hidden uh and entangled

08:24.350 --> 08:26.517
in in high dimensional space , right .

08:26.517 --> 08:28.739
So we , we want to build a transform to

08:28.739 --> 08:30.799
disentangle and , and , and put that

08:30.809 --> 08:33.520
into a new space . Uh But a general

08:33.530 --> 08:35.919
goal um more specific than just to

08:35.929 --> 08:38.040
bring the structure more explicit . A

08:38.040 --> 08:40.096
general goal is to transform the rot

08:40.096 --> 08:42.330
into a new space such that the similar

08:42.700 --> 08:46.080
things are placed closer . And

08:46.090 --> 08:47.923
meanwhile , the new space is not

08:47.923 --> 08:49.646
collapsed , right ? So you can

08:49.646 --> 08:51.868
absolutely put everything close to each

08:51.868 --> 08:54.090
other . But uh it's at the same point ,

08:54.090 --> 08:56.201
right . So it's trivial . So we don't

08:56.201 --> 08:58.257
want to be collapsed , right ? So we

08:58.257 --> 09:00.479
still want to maintain the volume . But

09:00.479 --> 09:02.534
in uh in that new space , similar uh

09:02.534 --> 09:04.757
similarities really get reflected . And

09:04.757 --> 09:06.534
the second question immediately

09:06.534 --> 09:08.590
following from that is that uh where

09:08.590 --> 09:10.590
does uh uh supervision come from or

09:10.590 --> 09:12.701
that similarity come from , there are

09:12.701 --> 09:15.849
really cla three classical ideas appear

09:15.859 --> 09:18.349
again and again . Um in , in the

09:18.359 --> 09:20.248
literature , the first one is the

09:20.248 --> 09:22.359
spatial co occurrence , right . So we

09:22.359 --> 09:24.359
look at this thought when if we are

09:24.359 --> 09:26.359
thinking about these uh each of the

09:26.359 --> 09:28.581
image patches , they really cour within

09:28.581 --> 09:31.070
this uh same con context , right ? So

09:31.080 --> 09:33.469
when they cour it shows some similarity

09:33.479 --> 09:35.989
between them um that spatial core

09:36.000 --> 09:38.167
occurrence . And the second one is the

09:38.167 --> 09:40.280
temporal core occurrence when we have

09:40.289 --> 09:42.511
the natural videos , right . So each of

09:42.511 --> 09:44.640
the net uh video frames , they follow

09:44.650 --> 09:46.872
each other , right . So , and when they

09:46.872 --> 09:49.039
happen in the same context , they have

09:49.039 --> 09:51.150
some similarity . And the third ideas

09:51.150 --> 09:53.400
come from uh mathematics is really

09:53.409 --> 09:55.679
about in the high dimensional space if

09:55.690 --> 09:58.530
we have each of the data point , all of

09:58.539 --> 10:00.650
them , right ? So in high dimensional

10:00.650 --> 10:02.761
space and for each of the points , we

10:02.761 --> 10:04.872
can find a neighborhood . And when we

10:04.872 --> 10:07.179
put that neighborhood , um uh all these

10:07.190 --> 10:09.301
together for each of the points , the

10:09.301 --> 10:11.468
neighborhood uh and their neighborhood

10:11.468 --> 10:13.780
together will show some of the geometry

10:14.020 --> 10:16.280
in the data space , right . So , and by

10:16.289 --> 10:18.880
using that , that's a Euclidean local ,

10:18.890 --> 10:20.960
very local Euclidean um uh uh

10:20.969 --> 10:23.090
neighborhood that can show you

10:23.099 --> 10:25.155
something about the data structure .

10:25.859 --> 10:28.320
These are the three major uh classical

10:28.330 --> 10:30.052
ideas appear again and again ,

10:30.052 --> 10:32.330
computing neuroscience , deep learning .

10:32.330 --> 10:34.719
And we basically uh leverage these uh

10:34.729 --> 10:38.080
insect and they also uh generalized to

10:38.090 --> 10:40.210
other modality . When it comes to

10:40.219 --> 10:42.386
natural language , we can either treat

10:42.386 --> 10:44.489
that as a spatial co occurrence or a

10:44.500 --> 10:48.070
temporal co occurrence . Um All right .

10:48.080 --> 10:51.159
So today , uh my main point would be

10:51.169 --> 10:53.630
that on one hand , we want to build our

10:53.679 --> 10:56.510
NW representation transform derived

10:56.520 --> 10:58.809
from neural and sta principle , right .

10:58.820 --> 11:01.270
So that's um um one direction , right .

11:01.280 --> 11:03.580
So very low level and uh and , and ,

11:03.590 --> 11:05.923
and we want to build it up , right ? So ,

11:05.923 --> 11:08.146
and to accomplish our goal , right ? So

11:08.146 --> 11:10.200
just to define and representation

11:10.210 --> 11:12.488
learning and to reflect the similarity .

11:12.950 --> 11:15.006
On the other hand , there has been a

11:15.010 --> 11:17.780
lot of the deep learning and deep self

11:18.299 --> 11:20.466
learning method , right ? We want to ,

11:20.466 --> 11:22.632
but we want to understand them right ,

11:22.632 --> 11:24.940
by simplification and reductionism and

11:24.950 --> 11:28.090
unification of them . And any uh at the

11:28.099 --> 11:30.043
end , I will show their surprising

11:30.043 --> 11:32.210
convergence and they share , it turns

11:32.219 --> 11:34.052
out they share the same learning

11:34.052 --> 11:36.275
objective and their engineering gap can

11:36.275 --> 11:38.219
also be largely closed , not fully

11:38.219 --> 11:40.940
closed yet , but largely closed . All

11:40.950 --> 11:43.229
right . So uh this is really , today's

11:43.239 --> 11:46.690
talk is really , oh , it's gone . Hey ,

11:46.700 --> 11:49.609
hey , I , I didn't know if it would be

11:49.619 --> 11:51.508
OK for me to interrupt you with a

11:51.508 --> 11:53.619
question or not . Uh uh Yeah . Yeah .

11:53.619 --> 11:55.675
Please go ahead . Oh OK . OK . Uh uh

11:55.675 --> 11:57.786
It's uh I'm sorry to cut you off here

11:57.786 --> 11:59.952
at the end of your , your this slide .

11:59.952 --> 12:02.008
But um on the previous , you had the

12:02.008 --> 12:04.510
three the three sources of similarity

12:04.520 --> 12:06.687
there and , and the way that you , you

12:06.687 --> 12:09.619
talk about the uh the similarity in uh

12:09.630 --> 12:11.741
the , the bottom example , you know ,

12:11.741 --> 12:13.741
person to spatially to all of those

12:13.741 --> 12:15.797
other , to those other words . Uh Do

12:15.797 --> 12:17.852
you think that that there's a way to

12:17.852 --> 12:19.963
map ? Uh and maybe I'm thinking about

12:19.963 --> 12:21.519
it wrong s semantic kind of

12:21.519 --> 12:23.574
similarities . So I'm thinking about

12:23.574 --> 12:25.797
some of these like concept net and some

12:25.797 --> 12:25.409
of these systems where people have

12:25.419 --> 12:27.252
intentionally tried to tease out

12:27.252 --> 12:29.197
semantic similarities from uh from

12:29.197 --> 12:31.308
crowdsourcing kinds of things or , or

12:31.308 --> 12:33.252
even like a RLHF and some of these

12:33.252 --> 12:35.308
tools for , for language models . Do

12:35.308 --> 12:37.530
you think there's a , a way that , that

12:37.530 --> 12:39.863
maps into these categories ? Um uh Yeah ,

12:39.863 --> 12:42.140
I think semen the semantic similarity

12:42.150 --> 12:44.840
is a , a quite high level in the sense

12:44.849 --> 12:46.960
that a lot of these could be uh feels

12:46.960 --> 12:49.071
low level , right ? So , and , but it

12:49.071 --> 12:50.989
also has a semantic and synthetic

12:51.460 --> 12:53.630
synthetic when it comes to natural

12:53.640 --> 12:55.640
language , right ? So for example ,

12:55.640 --> 12:58.739
this person , right . So it um um it

12:58.750 --> 13:02.020
will co uh cour with many of the words

13:02.030 --> 13:04.390
that reflect the different semantic and

13:04.400 --> 13:07.919
the synthetic um uh uh uh uh structure

13:07.929 --> 13:09.762
of this uh particular word , for

13:09.762 --> 13:12.369
example , this met , right . So , and a

13:12.570 --> 13:14.799
uh right . So who , right . So that

13:14.809 --> 13:17.159
reflects what this um word actually

13:17.169 --> 13:20.250
mean ? Right . So , and , and this is

13:20.260 --> 13:23.229
uh still try to build white box models

13:23.239 --> 13:25.719
that can um um form a language

13:25.729 --> 13:27.896
representation still in the very early

13:27.896 --> 13:30.469
stage , right ? So I , and presenting

13:30.479 --> 13:32.659
today will probably shed some light on

13:32.750 --> 13:35.260
how to understand the word embeddings ,

13:35.309 --> 13:37.531
right ? So which is a lot , low level ,

13:37.531 --> 13:40.789
a lot more low uh well , uh lower level

13:40.809 --> 13:42.862
uh compared to the large language

13:42.872 --> 13:44.816
models , right ? So we still don't

13:44.816 --> 13:46.983
fully understand that . But um I think

13:46.983 --> 13:49.094
the general idea will generalize it ,

13:49.094 --> 13:51.094
right . So we if we build on top of

13:51.094 --> 13:52.983
that and continue to develop that

13:52.983 --> 13:55.413
potentially , we can um have much more

13:55.422 --> 13:57.700
understanding about our language model .

13:57.783 --> 13:59.894
And in fact , some of these ideas can

13:59.894 --> 14:02.403
be applied to uh explain what's going

14:02.413 --> 14:04.524
on in large language models , right ,

14:04.524 --> 14:06.935
to open it up , right . So Anthro Top

14:07.005 --> 14:09.596
has uh follow up uh our work in

14:09.606 --> 14:12.315
visualizing transformers . And we did

14:12.325 --> 14:14.616
some of the earliest uh uh um

14:14.625 --> 14:16.776
mechanistic explanation for large

14:16.786 --> 14:18.906
language models by using dictionary

14:18.916 --> 14:21.027
learning to open it up , right ? Show

14:21.027 --> 14:23.255
that in the early stage that you can do

14:23.265 --> 14:26.585
disambiguation , right . So left , left

14:26.596 --> 14:28.763
the room turn left , right . So that's

14:28.763 --> 14:30.818
very different , right ? In the very

14:30.818 --> 14:33.096
early stage language model can do that .

14:33.096 --> 14:35.263
And then uh in the middle layers , you

14:35.263 --> 14:37.485
can form new in new formation , right .

14:37.485 --> 14:39.707
So you can find that OK , for example ,

14:39.707 --> 14:41.985
uh from uh the the unit change , right .

14:41.985 --> 14:44.210
So from , from C Celsius to uh front

14:44.330 --> 14:48.215
head or from uh centimeters to me are

14:48.224 --> 14:50.391
uh you know , the bad transform in the

14:50.391 --> 14:52.557
high level , you can do the repetitive

14:52.557 --> 14:54.557
pattern detections , you know , all

14:54.557 --> 14:56.724
this can be um you can use some of the

14:56.724 --> 14:59.275
ideas in this talk to open up the

14:59.284 --> 15:01.724
language model , right ? So yeah , but

15:01.734 --> 15:05.380
um still in the early stage , I hope

15:05.390 --> 15:07.668
uh uh that , that answer your question .

15:07.668 --> 15:09.834
Oh yeah , yeah , thank you . Thank you

15:09.834 --> 15:12.112
very much . Oh OK , cool , cool , cool .

15:12.112 --> 15:14.001
Um So today's talk is a lot about

15:14.001 --> 15:16.140
science and math . Uh where the , the

15:16.150 --> 15:18.150
the central question is , how do we

15:18.150 --> 15:20.483
understand the signal space and learn a ,

15:20.483 --> 15:23.669
learn or build a white box uh model

15:23.679 --> 15:25.989
accordingly ? Right . So , and on the

15:26.000 --> 15:28.167
other hand , my group also engaged the

15:28.167 --> 15:31.549
other very uh uh this is a microscopic

15:31.559 --> 15:33.989
and the other macroscopic uh idea is

15:34.000 --> 15:36.659
about engineering . How can we build um

15:36.669 --> 15:38.613
um sample efficient world models ,

15:38.613 --> 15:40.725
right . So that's um you know that uh

15:40.725 --> 15:42.750
but today's uh is about more about

15:42.760 --> 15:45.619
science and mathematics . Um OK . So ,

15:45.630 --> 15:48.359
and was representation transform

15:48.369 --> 15:51.469
derived from neural an Staal principles ,

15:51.479 --> 15:53.701
right ? So that's uh and , and , and uh

15:53.701 --> 15:55.701
the first principle shall come from

15:55.701 --> 15:57.923
computational neuroscience in the in uh

15:57.923 --> 15:59.868
compu uh neuroscientists have long

15:59.869 --> 16:02.455
observed that uh from LGN , right ,

16:02.465 --> 16:05.474
which is really relay from the uh the

16:05.484 --> 16:08.515
retina to the visual cortex from LGN to

16:08.525 --> 16:11.465
the early visual cortex . V one that uh

16:11.474 --> 16:13.755
dimensionality has been expanded by

16:13.765 --> 16:17.424
1000 hundreds to tho 1000 times , right ?

16:17.434 --> 16:21.280
So , and an early success and

16:21.289 --> 16:23.500
um um in computational neuroscience is

16:23.510 --> 16:25.599
that we , we try to understand why

16:25.609 --> 16:27.776
that's uh that's the case , right . So

16:27.776 --> 16:29.776
meanwhile , the the activity of all

16:29.776 --> 16:32.380
these neurons are super sparse compared

16:32.390 --> 16:34.880
to LGN , right . So , and , and , and

16:34.890 --> 16:37.210
early success in enterprise

16:37.219 --> 16:39.859
representation learning is uh uh sparse

16:39.869 --> 16:41.647
coding or in general dictionary

16:41.647 --> 16:43.849
learning is that the idea is really

16:43.859 --> 16:46.489
that simple , given an image , we

16:46.500 --> 16:48.770
assume that we can learn a set of

16:48.780 --> 16:50.950
common features which can be reused

16:50.960 --> 16:53.830
again and again and in a linear fashion ,

16:53.840 --> 16:56.650
right ? So that's a neural activity and

16:56.659 --> 16:59.390
uh that demonstrate uh determines how

16:59.400 --> 17:01.549
strongly a feature is used in

17:01.559 --> 17:03.615
reconstructing this particular image

17:03.615 --> 17:06.709
plus uh some noise , right . So that's

17:06.719 --> 17:09.010
um try to give each of the image , we

17:09.020 --> 17:11.709
try to linearly and sparsely decompose

17:11.719 --> 17:15.030
this uh image vector into a set of

17:15.040 --> 17:17.151
features sparse combination of that .

17:17.310 --> 17:21.180
Um once we train our um infer

17:21.189 --> 17:24.390
this activity and also learn these

17:24.400 --> 17:27.170
common features that we will find that

17:27.180 --> 17:29.291
these uh dictionary elements start to

17:29.291 --> 17:31.709
converge to things like this G boyish

17:31.719 --> 17:34.359
filters , right ? So that's uh that has

17:34.369 --> 17:37.790
a uh uh uh a shocking correspondence to

17:37.800 --> 17:41.390
the the um uh to the receptive field in

17:41.400 --> 17:44.859
the visual cortex or early cortex . And

17:44.869 --> 17:47.609
that really learns the structure in the

17:47.619 --> 17:50.069
early stage of the images , right . So

17:50.079 --> 17:51.968
low level statistics , right . So

17:51.968 --> 17:54.246
because a lot of these filters and why ,

17:54.246 --> 17:56.357
why should we expect them ? Because a

17:56.357 --> 17:58.579
lot of the times the natural images are

17:58.579 --> 18:01.349
uh contains a lot of these edge like

18:01.640 --> 18:03.696
structures , right ? So , and it's a

18:03.696 --> 18:06.050
it's a good idea to , to learn these

18:06.060 --> 18:08.282
edge like filters , right . So then you

18:08.282 --> 18:10.699
can assemble them and come uh uh to

18:10.709 --> 18:12.876
reconstruct the image more efficiently

18:12.876 --> 18:14.876
and sparsely , right . So that's uh

18:14.876 --> 18:16.820
here's the optimization , it's the

18:16.820 --> 18:18.931
iterative optimization , right ? It's

18:18.931 --> 18:20.987
given the image we want to infer the

18:20.987 --> 18:23.153
sparse coefficient . And then after we

18:23.153 --> 18:25.042
have that , we want to update our

18:25.042 --> 18:27.265
dictionary or our features . And you do

18:27.265 --> 18:30.055
this , you know um iterative um convex

18:30.064 --> 18:32.231
optimization overall , it's not convex

18:32.231 --> 18:34.008
but iteratively that's a convex

18:34.008 --> 18:36.925
optimization . Um And you learn this uh

18:36.935 --> 18:39.589
dictionary . Um Yeah , let's use that

18:39.599 --> 18:41.710
right . So that's uh you know uh we ,

18:41.710 --> 18:43.821
we already have some representation .

18:43.821 --> 18:45.877
And let's see now given the original

18:45.877 --> 18:48.099
natural video , let's use the sparse uh

18:48.099 --> 18:50.829
dictionary to decompose our our signal

18:50.839 --> 18:54.199
into uh sparse coefficient space . And

18:54.209 --> 18:56.265
uh as you can see from the left hand

18:56.265 --> 18:58.376
side to the right hand side , you can

18:58.376 --> 19:00.487
see that the reconstruction is good ,

19:00.487 --> 19:02.653
right . So , and uh but meanwhile , if

19:02.653 --> 19:04.653
you look at the the response of the

19:04.653 --> 19:07.859
sparse coefficient is very random ish

19:07.869 --> 19:10.140
looking right . So if we revisit our

19:10.150 --> 19:13.099
ideas that uh we want to put si similar

19:13.109 --> 19:16.550
things uh in the latent space close and

19:16.560 --> 19:18.560
since that we didn't accomplish our

19:18.560 --> 19:20.616
goal , right . So the similar things

19:20.616 --> 19:22.616
seems not close to each other . But

19:22.616 --> 19:24.969
even more far apart , right . So uh um

19:25.829 --> 19:28.250
so , so what's missing here and we

19:28.260 --> 19:30.520
shall look at uh the second principle

19:30.530 --> 19:33.430
from statistical uh principle , right ?

19:33.439 --> 19:35.439
So it's called manifold learning or

19:35.439 --> 19:38.119
from mathematics . The idea is that in

19:38.130 --> 19:41.530
a high dimensional space , we imagine

19:41.540 --> 19:44.170
our data lives in a low-dimensional

19:44.180 --> 19:47.079
nonlinear uh subspace , right ? So in

19:47.089 --> 19:49.145
this case , it's a three dimensional

19:49.145 --> 19:51.239
space and our data reside on a two

19:51.290 --> 19:53.890
dimensional manifold , right ? So like

19:53.900 --> 19:56.410
this withdraw , right . So uh that in

19:56.420 --> 19:58.642
high dimensional , how , how do we find

19:58.642 --> 20:01.739
this uh uh uh uh structure is that for

20:01.750 --> 20:03.861
each of the data points , we can look

20:03.861 --> 20:07.439
around , find its neighbors and uh and

20:07.449 --> 20:09.719
only locally , right ? So if you , you

20:09.729 --> 20:11.989
know , if you generalize this

20:12.000 --> 20:14.979
neighborhood to a larger um larger um

20:14.989 --> 20:17.100
space , right ? So you , you start to

20:17.100 --> 20:20.010
get short circuit effect but um for

20:20.020 --> 20:22.187
each of the data points , you can find

20:22.187 --> 20:24.298
this neighborhood together . You will

20:24.298 --> 20:26.900
start to find um the , the the global

20:26.910 --> 20:28.959
geometry of your data , right ? So

20:28.969 --> 20:31.080
that's a , that's a basic idea almost

20:31.080 --> 20:33.358
like how we measured the earth , right ?

20:33.358 --> 20:35.302
So locally it's flat , flat , flat

20:35.302 --> 20:37.636
together when we saw them together . OK .

20:37.636 --> 20:39.913
That's , that's and , and uh uh sphere ,

20:39.913 --> 20:42.010
right . So , and then once we have

20:42.020 --> 20:44.187
these all these neighborhoods , we can

20:44.187 --> 20:46.298
aim to reduce the dimensionality to a

20:46.298 --> 20:47.853
low ier space such that the

20:47.853 --> 20:49.742
neighborhood structure are um are

20:49.742 --> 20:52.400
respected . So that's the idea it's

20:52.410 --> 20:55.050
very relevant in model uh in modeling

20:55.060 --> 20:57.699
natural transform . Even the simplest

20:57.709 --> 20:59.859
natural transformations in in in the

20:59.869 --> 21:02.091
image domain would follow some of these

21:02.091 --> 21:04.147
manifold structure . For example , I

21:04.147 --> 21:06.339
have this letter a image letter A from

21:06.349 --> 21:08.516
the left hand side to translate to the

21:08.516 --> 21:10.627
right hand side , right . So that's a

21:10.627 --> 21:12.460
very , very nonlinear trajectory

21:12.460 --> 21:14.627
because it's if it's linear and we can

21:14.627 --> 21:16.793
do linear interpolation and that's the

21:16.793 --> 21:18.960
linear interpolation get us right . So

21:18.960 --> 21:21.016
it's the superposition right . So on

21:21.016 --> 21:22.793
the other hand , if we have the

21:22.793 --> 21:24.571
manifold and we can do manifold

21:24.571 --> 21:26.738
interpolation get in between , right .

21:26.738 --> 21:28.905
So that's uh um uh that , that's a one

21:28.905 --> 21:30.960
natural transformations but there in

21:30.960 --> 21:33.016
real world there could be many other

21:33.016 --> 21:35.127
transformation like 3d rotation light

21:35.127 --> 21:37.182
shading change uh right . So that uh

21:37.182 --> 21:38.960
this all follows uh manifold um

21:38.960 --> 21:42.890
structures . And um here's the uh idea ,

21:42.900 --> 21:44.900
right . So how do we one particular

21:44.900 --> 21:46.956
example ? My favorite example , even

21:46.956 --> 21:48.733
though there are many different

21:48.733 --> 21:50.900
formulations , they are really all the

21:50.900 --> 21:53.122
same , right . So they , they have here

21:53.122 --> 21:55.289
and there . But uh the central idea is

21:55.289 --> 21:57.511
uh is um all the same . Let's um use um

21:57.511 --> 21:59.678
this uh local linear embedding , which

21:59.678 --> 22:01.844
is the classical one as a very elegant

22:01.844 --> 22:04.050
formulation um for each of the data

22:04.060 --> 22:07.329
point X I , right . So we want to find

22:07.339 --> 22:09.599
its neighborhood and local euan

22:10.000 --> 22:12.056
neighborhood , right . So and then ,

22:12.770 --> 22:14.890
and we can use that this neighborhood

22:14.900 --> 22:17.949
to do a linear interpolation ,

22:18.579 --> 22:22.390
right . So here is a data point uh data

22:22.400 --> 22:24.510
points . And uh if we use the whole

22:24.520 --> 22:27.229
data set as our dictionary , and we can

22:27.770 --> 22:30.630
particularly select its neighborhood to

22:30.640 --> 22:32.500
do this linear interpolation to

22:32.510 --> 22:35.140
reconstruct this data point , right .

22:35.160 --> 22:37.260
So our dictionary is huge , right ?

22:37.270 --> 22:39.326
Because we use the whole data set as

22:39.326 --> 22:41.548
our dictionary , right ? So our feature

22:41.548 --> 22:44.750
um and and that's gonna be a very ,

22:44.760 --> 22:47.500
very sparse and linear reconstruction

22:47.510 --> 22:49.732
for each of the data point , right . So

22:49.732 --> 22:51.954
because you know , among the whole data

22:51.954 --> 22:54.121
set , we , we select that neighborhood

22:54.121 --> 22:56.232
only a few data point amount millions

22:56.232 --> 22:58.343
of points , right . And they , they ,

22:58.343 --> 23:00.454
they sum to one , right . So that's a

23:00.454 --> 23:03.520
very sparse linear reconstruction . And

23:03.530 --> 23:05.586
um and together for each of the data

23:05.586 --> 23:08.060
point we can do so and we can rewrite

23:08.069 --> 23:11.130
this objective and , and after we um

23:11.140 --> 23:13.349
rewrite everything right . So we'll

23:13.359 --> 23:15.470
have that uh you know , that's our uh

23:15.470 --> 23:17.692
uh uh matrix form , our data set , each

23:17.699 --> 23:20.050
of the column means uh it's a uh data

23:20.060 --> 23:23.550
vector . And once we manipulate that ,

23:23.560 --> 23:26.685
we reach a , a optimization , we want

23:26.694 --> 23:29.984
to map that to a uh a low dimensional

23:29.994 --> 23:32.844
space , right . So by projecting them

23:32.854 --> 23:34.685
into a low dimensional space ,

23:34.694 --> 23:37.655
basically embed our data new data

23:37.665 --> 23:40.194
points into a low dimensional space .

23:40.204 --> 23:42.474
Uh such that this local linear

23:42.484 --> 23:45.525
embedding structure is reflected ,

23:46.349 --> 23:47.849
right ? We want to do this

23:47.849 --> 23:49.960
dimensionality reduction such that uh

23:49.960 --> 23:51.905
this uh local linear interpolation

23:51.905 --> 23:54.400
relationship is uh respected , right .

23:54.410 --> 23:56.569
So , um and we can rewrite this

23:56.579 --> 23:59.050
optimization uh with some manipulation

23:59.060 --> 24:00.838
and then we reach this spectral

24:00.838 --> 24:03.060
decomposition problem , right . So it's

24:03.060 --> 24:05.650
really that maintain this linear uh

24:05.660 --> 24:08.670
local linear embedding and maintain ,

24:08.680 --> 24:11.040
make sure that the transform is not

24:11.050 --> 24:13.209
collapsed , right . So , and , and ,

24:13.219 --> 24:15.770
and , and we want to solve P right ,

24:15.780 --> 24:18.255
once we solve for P and from a high

24:18.295 --> 24:20.462
dimensional space to a low dimensional

24:20.462 --> 24:22.295
space , and yet the local linear

24:22.295 --> 24:24.128
interpolation relationship is uh

24:24.128 --> 24:26.964
reflected , respected . And we should

24:26.974 --> 24:30.724
discover rediscover this um geometry in

24:30.734 --> 24:33.555
the our target embedding space , right .

24:33.564 --> 24:37.219
So that's the basic idea . Um if we put

24:37.229 --> 24:40.189
these sparse coding uh uh mathematical

24:40.199 --> 24:42.500
formulation and manifold learning

24:42.510 --> 24:44.869
formulation together back to uh back to

24:44.880 --> 24:47.750
back , we should find this surprise um

24:47.760 --> 24:50.930
uh similarity in their early uh the the

24:50.939 --> 24:53.050
their formulation is that the both of

24:53.050 --> 24:55.106
for each of the data point , both of

24:55.106 --> 24:56.606
them use a line and sparse

24:56.606 --> 24:58.717
decomposition . And for sports coding

24:58.717 --> 25:00.939
something is missing here because um as

25:00.939 --> 25:03.106
manifold learning has the second stage

25:03.106 --> 25:05.050
is that to do this low dimensional

25:05.050 --> 25:08.229
embedding such that re respect the the

25:08.239 --> 25:10.406
structure . But on the other hand , if

25:10.406 --> 25:12.517
you look at sparse coding , it's just

25:12.517 --> 25:14.628
to stop there , right . So , and very

25:14.628 --> 25:16.850
simple idea is that can we uh introduce

25:16.850 --> 25:19.199
the second stage embedding idea to

25:19.209 --> 25:21.319
sparse coding and to complete it ,

25:21.329 --> 25:23.420
right . So because that's where you

25:23.430 --> 25:26.020
really get to reflect the similarity in

25:26.030 --> 25:28.197
manifold learning , right . So now you

25:28.197 --> 25:30.630
know , if we combine that , what does

25:30.640 --> 25:32.900
that mean , right . So here is a very

25:32.910 --> 25:36.349
low dimensional one dimensional uh

25:36.359 --> 25:38.599
manifold in a three dimensional space

25:38.609 --> 25:40.665
on a sphere , right . So , and let's

25:40.665 --> 25:42.831
initialize our sparse coding , right .

25:42.831 --> 25:45.660
So , and then now let's use our sparse

25:45.819 --> 25:48.209
coding to learn that after you learn

25:48.219 --> 25:50.439
that it should start to reflect the

25:50.449 --> 25:52.709
local structures of your data and the

25:52.719 --> 25:55.599
features , right . So um and in this

25:55.609 --> 25:59.060
case , it will tile the manifold and um

25:59.069 --> 26:01.459
and , and if we stop here and here's

26:01.469 --> 26:04.430
the trajectory when our point is

26:04.439 --> 26:07.739
traverse a natural trajectory along the

26:07.750 --> 26:10.839
manifold , as you can see this alpha or

26:10.849 --> 26:13.739
our sparse code is relatively

26:13.750 --> 26:16.060
unstructured , seems random just as we

26:16.069 --> 26:18.510
saw in that natural video . But let's

26:18.520 --> 26:21.770
say if we use this manifold embedding

26:21.780 --> 26:24.150
into a new space , right ? So we should

26:24.160 --> 26:26.540
suddenly discover that OK , that's

26:26.550 --> 26:28.383
actually following a very linear

26:28.383 --> 26:31.339
trajectory in this um embedding space .

26:31.949 --> 26:34.171
That that's the basic idea , right . So

26:34.171 --> 26:37.250
we need this embedding uh step to

26:37.260 --> 26:40.400
complete this dictionary learning . And

26:41.079 --> 26:43.246
that leads to the formulation of spars

26:43.246 --> 26:45.989
manifold transform , right . So um a

26:46.000 --> 26:49.650
better idea the word may be a sparse

26:49.660 --> 26:52.540
spectral transform but um uh um

26:52.550 --> 26:55.609
manifold , it's um it's not bad , right ?

26:55.619 --> 26:58.880
So uh given the signal , we first uh

26:58.890 --> 27:01.319
use sparse coding to do a sparse and

27:01.329 --> 27:04.060
linear decomposition , right . So after

27:04.069 --> 27:07.469
that , we seek to embed our smart

27:08.030 --> 27:11.060
code into a new space , given our

27:11.069 --> 27:13.280
insight that you know these dictionary

27:13.290 --> 27:16.030
uh elements will collaborate with each

27:16.040 --> 27:18.439
other in order to model something in

27:18.449 --> 27:20.984
between a particular example is that if

27:20.994 --> 27:23.272
you think about these uh Gabor filters ,

27:23.272 --> 27:25.105
these Gabor filters and slightly

27:25.105 --> 27:26.938
shifted Gabor filters , they are

27:26.938 --> 27:28.883
neighbors , right ? So if you have

27:28.883 --> 27:30.883
something in between these two will

27:30.883 --> 27:33.435
work together and , and do this local

27:33.444 --> 27:35.900
linear interpolation very , very

27:35.910 --> 27:38.540
locally , right ? So , and , and if you

27:38.599 --> 27:42.430
um uh um project that into the

27:42.439 --> 27:44.780
embedding space , that should give you

27:44.790 --> 27:48.560
a more uh um uh give you

27:48.569 --> 27:52.439
more reflect uh uh um it will reflect

27:52.449 --> 27:54.510
the similarity better because uh you

27:54.520 --> 27:56.742
know , if you don't do this embedding ,

27:56.742 --> 27:58.798
you will find OK , this one is going

27:58.798 --> 28:00.798
down and the other one's going up ,

28:00.798 --> 28:02.853
right ? So you , you know , it feels

28:02.853 --> 28:04.964
like random noise . But if you know ,

28:04.964 --> 28:07.076
you know their neighbors , right ? So

28:07.076 --> 28:07.000
you put that in there as neighbors in

28:07.010 --> 28:09.189
the embedding space and that start to

28:09.199 --> 28:12.609
give you uh back the similarity . And

28:12.619 --> 28:14.675
uh I don't know how many of you have

28:14.675 --> 28:16.508
heard this um uh low dimensional

28:16.508 --> 28:19.930
analysis uh of the natural image um uh

28:19.939 --> 28:22.369
high contrast patches by uh Gunner Ken

28:22.729 --> 28:24.896
at Stanford , right ? So , and there's

28:24.896 --> 28:28.099
a very interesting um thing is that it

28:28.109 --> 28:30.760
in uh date back to the analysis that De

28:30.780 --> 28:33.329
Man Ford has done uh a long time ago

28:33.339 --> 28:36.060
with Tesson Lee . Um is that for high

28:36.069 --> 28:38.069
contrast , the natural images , you

28:38.069 --> 28:41.020
will um you can do some preprocessing

28:41.030 --> 28:43.020
and very low level three by three

28:43.030 --> 28:44.808
natural images , right ? So and

28:44.808 --> 28:46.752
eventually we will conclude , OK ,

28:46.752 --> 28:48.989
that's a line bottle , right ? So you

28:49.000 --> 28:51.810
analyze this topological analysis .

28:51.819 --> 28:53.541
Really ? I I think that's very

28:53.541 --> 28:55.708
interesting because um you know , even

28:55.708 --> 28:57.652
though now we can build the neural

28:57.652 --> 28:59.652
networks that can recognize a cat ,

28:59.652 --> 29:01.486
right ? So we still have no clue

29:01.486 --> 29:03.800
exactly what a three by three natural

29:03.810 --> 29:06.109
image patch space look like we have

29:06.119 --> 29:07.952
some clue , right ? So but not a

29:07.952 --> 29:10.063
complete clue , right ? So one of the

29:10.063 --> 29:12.589
partial clue is that if we focus on the

29:12.599 --> 29:15.349
high contrast natural images and we'll

29:15.359 --> 29:18.569
use topological analysis and uh use

29:18.579 --> 29:21.229
many numbers to analyze those holes .

29:21.239 --> 29:23.572
You will find that OK . That's a bottle .

29:23.572 --> 29:25.350
Cli bottle is a two dimensional

29:25.350 --> 29:27.814
manifold which can only exac uh reside

29:27.824 --> 29:29.991
in , in dimension higher than the four

29:29.991 --> 29:32.213
dimensional space , right . So and here

29:32.213 --> 29:34.435
is the explanation why that should be a

29:34.435 --> 29:36.380
client bo in high contrast natural

29:36.380 --> 29:38.380
images if we think about this Gabor

29:38.380 --> 29:40.546
filter space , right . So , and II I ,

29:40.546 --> 29:44.015
if you look at these allow the rotation ,

29:44.349 --> 29:46.682
which is natural transformation , right ?

29:46.682 --> 29:48.905
So , and also allows the phase change ,

29:48.905 --> 29:51.016
right ? So it's shifting , right ? So

29:51.016 --> 29:53.238
basically the shifting and you organize

29:53.238 --> 29:55.405
these dictionary element as a manifold

29:55.405 --> 29:57.516
that's basically a cli climb bottle .

29:57.516 --> 29:59.627
So you have my recording and then you

29:59.627 --> 30:01.738
can play with that after this meeting

30:01.738 --> 30:03.905
and you will find that actually that's

30:03.905 --> 30:06.016
the same topology as a climb bottle .

30:06.016 --> 30:08.071
And here are some mathematician here

30:08.071 --> 30:10.238
and and together with this uh trans uh

30:10.238 --> 30:12.460
translation and scaling , right ? So we

30:12.460 --> 30:14.405
really have a , a five dimensional

30:14.405 --> 30:16.516
manifold . In my original formulation

30:16.516 --> 30:18.349
of sparse manifold transform . I

30:18.349 --> 30:20.293
thought it it's a four dimensional

30:20.293 --> 30:22.238
manifold uh space . But um after I

30:22.238 --> 30:24.349
thought it more , it actually was a ,

30:24.349 --> 30:26.571
shouldn't be a mistake , it should be a

30:26.571 --> 30:28.800
five dimensional space . But um um

30:28.810 --> 30:31.000
intuition , I think it the the same .

30:31.290 --> 30:33.660
Um OK . So let's get back to the Spars

30:33.689 --> 30:35.979
Medical Transform formulation , right .

30:35.989 --> 30:38.719
So now we are uh we decompose our

30:38.729 --> 30:41.079
signal with sparse dictionary uh well

30:41.089 --> 30:43.089
with our dictionary in a sparse and

30:43.089 --> 30:45.219
linear fashion . And then we also

30:45.229 --> 30:47.760
imagine that our dictionary has some

30:47.770 --> 30:49.800
underlying manifold structure

30:49.810 --> 30:52.032
underlying them , right ? So we want to

30:52.032 --> 30:54.150
embed uh our dictionary element into

30:54.160 --> 30:57.719
the manifold space . And so after that ,

30:57.729 --> 31:00.520
we can use this embedding P to project

31:00.530 --> 31:02.530
our sparse signal into that space ,

31:03.060 --> 31:05.390
right . So here's the example that's

31:05.400 --> 31:07.622
our sparse coefficients , right . So uh

31:07.622 --> 31:09.844
when there's a natural transformation ,

31:09.844 --> 31:12.011
as you can see , imagine this as three

31:12.011 --> 31:13.900
edges , right ? So when there's a

31:13.900 --> 31:16.122
global transformations , right , so the

31:16.122 --> 31:17.956
decomposition really change very

31:17.956 --> 31:20.067
quickly because the support of sparse

31:20.067 --> 31:22.233
uh coefficient will also change very ,

31:22.233 --> 31:24.239
very quickly . And after this many

31:24.329 --> 31:26.400
embedding , we should expect that in

31:26.410 --> 31:28.300
the embedding space that's uh um

31:28.310 --> 31:30.290
linearized , right . So this

31:30.300 --> 31:34.050
formulation is very um uh uh uh uh is

31:34.060 --> 31:36.469
very much also inspired by slow feature

31:36.479 --> 31:38.780
analysis , right ? So in , in that

31:38.790 --> 31:40.810
sense you , you need this temporal

31:41.030 --> 31:43.449
similarity . But as I said in the early

31:43.459 --> 31:45.989
stage , I I'll get back to this um um

31:46.000 --> 31:49.880
uh the early uh uh similarity sources ,

31:49.890 --> 31:51.834
right . So you can generalize this

31:51.834 --> 31:54.959
formulation later on . And um this is a

31:55.140 --> 31:57.400
uh a linear interpolation in time

31:57.410 --> 31:59.521
domain , right . So which is a second

31:59.521 --> 32:02.609
order uh uh derivative , right . So ,

32:02.619 --> 32:04.839
and when we write rewrite this

32:04.849 --> 32:07.540
optimization , this embedding of our

32:07.550 --> 32:11.319
dictionary elements , Uh Again , we

32:11.329 --> 32:13.880
reach a a spectral decomposition

32:13.890 --> 32:16.079
problem right . When we solve that

32:16.089 --> 32:18.359
problem , we will solve for this P

32:18.369 --> 32:20.800
embedding our of our sparse uh

32:20.810 --> 32:22.866
dictionary element . And then we can

32:22.869 --> 32:25.339
use this P to do this transform . Uh

32:25.349 --> 32:27.770
let's try that right . So given a

32:27.780 --> 32:30.002
natural video , right ? So we can first

32:30.002 --> 32:32.002
use sparse coding to decompose that

32:32.002 --> 32:33.724
into a high dimensional sparse

32:33.724 --> 32:35.669
coefficient space . As you can see

32:35.669 --> 32:37.891
again , it's similarly random , right ?

32:37.891 --> 32:40.829
So , but let's try to embed our uh

32:40.839 --> 32:43.910
sparse coefficient into a new space ,

32:43.920 --> 32:46.060
right , do this uh dimension uh

32:46.069 --> 32:48.979
reduction and projection . And in this

32:48.989 --> 32:51.100
embedding space , as you can see that

32:51.100 --> 32:53.060
it's changing actually much slower

32:53.500 --> 32:55.680
space . And overall the transform can

32:55.689 --> 32:58.349
be approximately invertible . And now

32:58.359 --> 33:00.137
you can actually go back to the

33:00.137 --> 33:02.380
original space and you do lose a little

33:02.390 --> 33:05.199
bit uh uh detail . But again , it's um

33:05.209 --> 33:07.209
you , you keep the majority of them

33:07.209 --> 33:10.349
right . So um let's generalize the

33:10.359 --> 33:12.470
formulation because I I , what I just

33:12.470 --> 33:14.670
mentioned was use T derivative but

33:14.680 --> 33:17.189
again , you can use spatial derivative ,

33:17.199 --> 33:19.650
use that to uh because that can also

33:19.660 --> 33:22.540
tell you who are close to whom and uh

33:22.550 --> 33:24.383
what I just used the temporal co

33:24.383 --> 33:26.439
occurrence . And further you can use

33:26.439 --> 33:28.550
the Euclidean neighborhood right . So

33:28.550 --> 33:30.439
all these sources will give you a

33:30.439 --> 33:32.661
similar formulation . It's just how you

33:32.661 --> 33:34.494
define this interpolation in the

33:34.494 --> 33:36.550
embedding space , right . So whether

33:36.550 --> 33:38.883
that's a spatial temporal general right ,

33:38.883 --> 33:40.994
so or just a similarity which doesn't

33:40.994 --> 33:43.161
have to be first order derivative , it

33:43.161 --> 33:45.383
can be first order derivative , right .

33:45.383 --> 33:47.494
So , and when you each single time we

33:47.494 --> 33:49.772
write this um uh in such a formulation ,

33:49.772 --> 33:52.020
you and that optimization should be a

33:52.030 --> 33:54.363
spectral decomposition in order solving .

33:54.363 --> 33:56.510
Uh so that you solve that embedding .

33:56.930 --> 33:59.780
Um And you can use that to do uh word

33:59.790 --> 34:02.050
embedding uh formulation as well ,

34:02.060 --> 34:04.004
right . So that's uh that actually

34:04.004 --> 34:06.282
gives you a pretty good word embedding .

34:06.282 --> 34:08.760
And that can provide the word embedding

34:08.770 --> 34:12.100
explanation and explain what's going on

34:12.110 --> 34:14.620
in the word embedding . Umy of all

34:14.639 --> 34:18.000
these um All right . So that's the

34:18.010 --> 34:21.139
manifold embedding . And you can also

34:21.169 --> 34:24.250
uh uh uh has a strong relationship to

34:24.260 --> 34:26.316
graph embedding , right . So you can

34:26.316 --> 34:28.409
generalize this to , to the how to

34:28.419 --> 34:32.370
embed the graph nodes . Um OK . Kevin ,

34:32.770 --> 34:34.881
hey , um there's this question in the

34:34.881 --> 34:36.937
chat , is it OK ? If I read it , are

34:36.937 --> 34:39.310
there equivalencies here with cryo em

34:39.320 --> 34:42.860
analysis , cryo M analysis ? Well ,

34:42.870 --> 34:46.060
I'm analysis . I this word doesn't uh

34:46.070 --> 34:48.739
resonate with anything in my mind ? So

34:48.750 --> 34:51.159
I may need to take a look , right ? So

34:51.169 --> 34:53.409
uh please send me the uh the reference

34:53.419 --> 34:56.179
II , I apologize that I don't

34:56.189 --> 34:59.300
immediately uh know that word . Um Yeah ,

34:59.310 --> 35:01.729
sorry about that . Um I , I hope I can

35:01.739 --> 35:05.090
answer that . Uh But um uh uh so , so

35:05.100 --> 35:07.378
this uh uh graph embedding , you , you ,

35:07.378 --> 35:09.629
you , you has a similar idea , right ?

35:09.639 --> 35:11.583
So you , you , you for each of the

35:11.583 --> 35:13.806
notes , right ? So you can uh oh yeah ,

35:13.806 --> 35:16.028
yeah , that's my email address . So for

35:16.028 --> 35:18.083
the graph nodes and you can also use

35:18.083 --> 35:21.909
their AIC to define this um linearity ,

35:21.919 --> 35:24.086
right . So and then you can generalize

35:24.086 --> 35:26.820
this to linear uh general um for each

35:26.830 --> 35:28.886
of the node , right ? So you can use

35:28.886 --> 35:31.610
nearby ones to interpret that and then

35:31.620 --> 35:33.840
you can form that uh embedding

35:33.850 --> 35:36.128
formulation spectral embedding , right ?

35:36.128 --> 35:38.395
So that's um a close relationship . So

35:38.406 --> 35:40.684
that's a generalization work like that .

35:40.684 --> 35:43.426
And um there are also uh two

35:43.466 --> 35:46.535
perspective . I want to talk a little

35:46.545 --> 35:49.145
bit about this . Um I think that's very ,

35:49.156 --> 35:50.823
very important because when I

35:50.823 --> 35:53.016
originally formulated this problem , I

35:53.025 --> 35:55.736
I my mind was very continuous . I

35:55.746 --> 35:57.968
studied mathematics . My mind is very ,

35:57.968 --> 36:01.001
very uh continuous . I always better in

36:01.011 --> 36:03.122
in analysis and geometry than what I

36:03.132 --> 36:07.031
mean in algebra . But um and uh here is

36:07.041 --> 36:09.263
a meaningful point of view , right ? So

36:09.263 --> 36:11.961
if we look at this image , this um uh

36:11.971 --> 36:14.642
three image patches as you can see that

36:14.652 --> 36:17.511
they follow a very nice , that's a real

36:17.521 --> 36:19.799
image , right ? So , and , and you can ,

36:19.799 --> 36:22.021
you can see that it follows a very nice

36:22.021 --> 36:24.188
meaningful structure because this edge

36:24.188 --> 36:26.354
is moving here , moving here is highly

36:26.354 --> 36:28.929
me space that treats out a very nice uh

36:29.169 --> 36:32.489
um uh smooth curve even though uh it's

36:32.500 --> 36:34.860
highly non linear , but uh it's still a

36:34.870 --> 36:37.179
manifold structure . But if you look at

36:37.189 --> 36:39.429
this image , right . So , and let's say

36:39.439 --> 36:41.495
again , we want to put them close in

36:41.495 --> 36:43.606
the embedding space , right ? So this

36:43.606 --> 36:45.550
is a totally different perspective

36:45.550 --> 36:48.340
because this i this image patch is a

36:48.350 --> 36:52.219
dot le , right ? So it cour

36:52.540 --> 36:55.689
with this grassland cour with this uh

36:55.699 --> 36:58.870
thigh . Um co occur with some of these

36:58.879 --> 37:01.157
shadow and a another grassland , right ?

37:01.157 --> 37:02.990
So that context , this is a very

37:03.000 --> 37:05.639
noncontinuous , right ? But this co

37:05.739 --> 37:07.795
occurrence is also where the meaning

37:07.795 --> 37:09.795
come from because relationship come

37:09.795 --> 37:11.850
from uh well , the meaning come from

37:11.850 --> 37:14.310
relationship , right ? So a lag means

37:14.320 --> 37:17.370
lag is because it cour with all these

37:17.739 --> 37:21.010
uh uh uh uh oh yeah , some , some

37:21.020 --> 37:24.310
comments . OK . OK . So um so this l

37:24.570 --> 37:27.090
its meaning come from all the context

37:27.100 --> 37:29.322
of the cour , right ? So it or in other

37:29.322 --> 37:32.639
words , the red in one's head is , you

37:32.649 --> 37:34.538
know , has a lot of other meaning

37:34.538 --> 37:36.760
associated with the flower , you know ,

37:36.760 --> 37:38.871
o obvious it's not just that spectrum

37:38.871 --> 37:40.982
of the uh the the the light , right ?

37:40.982 --> 37:42.982
So that relationship determines the

37:42.982 --> 37:45.199
meaning and what I talked previously

37:45.209 --> 37:47.209
about manifold is only a continuous

37:47.209 --> 37:49.189
meaning , right ? So natural

37:49.199 --> 37:51.310
transformations , but this stochastic

37:51.310 --> 37:53.588
co occurrence is very , very important .

37:53.588 --> 37:56.530
Um uh I think , right ? So it , it ,

37:56.540 --> 37:59.010
it's a complementary perspective to

37:59.020 --> 38:02.770
share here um um uh and help us

38:02.780 --> 38:04.780
understand what's going on . But uh

38:04.780 --> 38:06.836
let's , let's also again , step into

38:06.836 --> 38:09.929
the uh detailed uh picture and see how

38:09.939 --> 38:11.995
this transform can help us solve our

38:11.995 --> 38:14.449
problem uh to accomplish our goal to

38:14.459 --> 38:16.681
transform the raw data into a new space

38:16.681 --> 38:18.681
such that similar things are placed

38:18.681 --> 38:20.681
closer and meanwhile not collapse ,

38:20.681 --> 38:22.909
right ? So I give you an example . Uh

38:22.919 --> 38:25.479
this example is a two entangled uh

38:25.489 --> 38:29.379
manifold as two spirals ,

38:29.389 --> 38:32.010
right ? So um entangled with each other

38:32.080 --> 38:35.270
in the original signal space , the

38:35.280 --> 38:37.590
neighborhood really is uh does not

38:37.600 --> 38:39.600
reflect the similarity because this

38:39.600 --> 38:42.689
this is a man uh spiral that uh

38:42.699 --> 38:45.500
these two points , right . So they are

38:45.510 --> 38:47.677
complete from two completely different

38:47.677 --> 38:49.899
manifold , right . So the distance does

38:49.899 --> 38:52.189
not actually mean much , right . So

38:52.199 --> 38:54.709
after our representation transform , we

38:54.719 --> 38:57.052
hope that when we talk about similarity ,

38:57.052 --> 38:59.275
we actually talk about this uh you know

38:59.275 --> 39:01.441
along the manifold , right . So that's

39:01.441 --> 39:03.386
our uh you know the the the goal ,

39:03.386 --> 39:05.659
right . So how we can accomplish our

39:05.669 --> 39:08.929
representation transform objective ?

39:09.479 --> 39:11.535
How do we do that ? Right . So first

39:11.535 --> 39:14.209
stage is that given this our data , we

39:14.219 --> 39:17.580
use the sparse coding to learn a set of

39:17.590 --> 39:19.923
features that can tell our data , right .

39:19.923 --> 39:22.090
So basically after we have learned our

39:22.090 --> 39:24.135
sparse dictionary , they , they

39:24.145 --> 39:27.574
basically tt our data space and build a

39:27.584 --> 39:29.473
support , right ? So that support

39:29.473 --> 39:31.640
doesn't tell you anything about who is

39:31.640 --> 39:33.975
close to whom . But at at the end , it ,

39:34.104 --> 39:36.784
it provides you the support so that you

39:36.794 --> 39:39.395
can draw your picture , right . So

39:39.405 --> 39:42.344
second , and we can use spectral um

39:42.354 --> 39:45.695
embedding to assign similar value to

39:45.705 --> 39:48.699
nearby features , right . So in this

39:48.709 --> 39:50.919
particular case , the the good idea is

39:50.929 --> 39:53.860
to assign the same value to all the uh

39:53.870 --> 39:55.899
nearby points on the , on the same

39:55.909 --> 39:58.830
manifold . And again , another constant

39:58.840 --> 40:01.007
value to the other points from another

40:01.007 --> 40:03.284
manifold , right . So that's the , the ,

40:03.284 --> 40:05.507
the the the most uh you know similar uh

40:05.510 --> 40:07.739
uh well , you give you the exactly the

40:07.750 --> 40:09.989
same number to all the points on the

40:10.000 --> 40:12.056
same manifold , right . So that's uh

40:12.056 --> 40:14.056
optimal piecewise constant function

40:14.056 --> 40:16.167
that it be defined on your manifold .

40:17.330 --> 40:19.552
And we can keep going right . So we can

40:19.552 --> 40:22.360
uh you know the second uh uh uh uh uh

40:22.370 --> 40:25.449
spectral embedding that assigns a , the

40:25.459 --> 40:27.820
similar point uh value to nearby point .

40:27.830 --> 40:30.052
But meanwhile , orthogonal to the first

40:30.052 --> 40:32.639
one um in in the function space is the

40:32.649 --> 40:34.760
another piecewise constant function .

40:34.760 --> 40:36.610
Now we run out of the piecewise

40:36.620 --> 40:39.439
constant function . And uh the next one

40:39.449 --> 40:41.727
would be the piecewise linear function ,

40:41.727 --> 40:43.870
right . So for each of these uh uh uh

40:43.919 --> 40:47.080
uh uh this embedding will assign

40:47.090 --> 40:49.379
linearly changing value to all the

40:49.389 --> 40:51.580
nearby points along one manifold ,

40:52.000 --> 40:54.056
right . So , and that's , that's the

40:54.056 --> 40:57.320
basic idea . It's um and once in the

40:57.330 --> 40:59.659
whole transform is basically coloring

40:59.669 --> 41:03.040
your space like this , right . So , um

41:03.050 --> 41:05.310
um uh again , that's uh from the data

41:05.320 --> 41:07.264
distribution , we force the use of

41:07.264 --> 41:09.580
sparse feature to learn a support to

41:09.590 --> 41:13.070
tile our uh uh uh data set , uh our

41:13.080 --> 41:15.191
data space , right . So , and then we

41:15.191 --> 41:17.358
can use low rank spectral embedding to

41:17.358 --> 41:20.340
color that space and establish that

41:20.350 --> 41:23.760
similarity . On top . Um I in some

41:23.810 --> 41:25.840
sense , you can see that as far as

41:25.850 --> 41:27.906
medical transform can be viewed as a

41:27.906 --> 41:30.370
data driven uh fouryear transform in

41:30.379 --> 41:32.629
fouryear transform , we have uh some

41:32.639 --> 41:34.583
support , right . So which is that

41:34.583 --> 41:36.639
could be a one dimensional Euclidean

41:36.639 --> 41:39.540
space or an image or uh um uh um uh uh

41:39.780 --> 41:43.370
tomography in three dimensional space ,

41:43.379 --> 41:45.889
right ? So you have our predefined

41:45.899 --> 41:48.360
support . But when you are given data ,

41:48.370 --> 41:51.040
right ? So can you define uh sparse uh

41:51.050 --> 41:53.459
can you define fourier transform on the

41:53.469 --> 41:55.525
data space ? Right . So , and sparse

41:55.620 --> 41:57.889
coding or diction learning provide you

41:57.899 --> 42:01.209
uh a support to treat out the um the

42:01.219 --> 42:03.600
the topology or the um data

42:03.610 --> 42:05.832
distribution , right . So , and then on

42:05.832 --> 42:07.666
top of that , you can build this

42:07.666 --> 42:09.770
spectral embedding right . So , and

42:09.780 --> 42:11.891
which would be a data driven employer

42:11.891 --> 42:14.590
transform in that sense . Um And let's

42:14.600 --> 42:16.711
use this representation to build some

42:16.711 --> 42:18.989
representation for natural images . And

42:19.000 --> 42:21.649
here's the example given the I of the

42:21.659 --> 42:25.399
image we first petrify this image

42:25.610 --> 42:27.721
into smaller spaces , right ? Because

42:27.721 --> 42:30.300
if you are in a high dimensional space ,

42:30.310 --> 42:33.229
right ? So that's um that's , that's

42:33.239 --> 42:36.840
not gonna be uh uh it's I I if you use

42:37.439 --> 42:39.439
uh dictionary to decompose that and

42:39.439 --> 42:42.149
that's still pretty hard . And if you ,

42:42.169 --> 42:44.280
you , in fact , each of the patches ,

42:44.280 --> 42:47.139
they got the uh decomposed by local

42:47.149 --> 42:49.316
dictionary elements , right ? So , and

42:49.316 --> 42:51.093
use the patch would uh would be

42:51.093 --> 42:53.610
simplifying that , that process . OK .

42:53.620 --> 42:56.070
Given each of the image and we can

42:56.080 --> 42:58.191
extract the patches . And for each of

42:58.191 --> 43:00.136
the patches , we can use sparse uh

43:00.136 --> 43:02.000
coefficient sparse function to

43:02.010 --> 43:04.179
decompose that . Right . So , and what

43:04.189 --> 43:06.411
now we are in a high dimensional sparse

43:06.411 --> 43:08.522
feature space and then we learn a low

43:08.522 --> 43:11.590
dimensional embedding uh to project

43:11.600 --> 43:13.211
this high dimensional sparse

43:13.211 --> 43:15.156
coefficient into a low dimensional

43:15.156 --> 43:17.322
space , right ? So and then we do some

43:17.322 --> 43:19.010
local pooling and aggregation

43:19.020 --> 43:21.429
normalization for that . And I just to

43:21.439 --> 43:24.530
do this one stage sparse uh sparse uh

43:24.540 --> 43:27.459
decomposition and then linear low rank

43:27.469 --> 43:31.149
linear uh embedding . And um

43:31.219 --> 43:33.840
we can build a very competitive

43:33.850 --> 43:36.017
transform . But before I show you that

43:36.159 --> 43:38.350
I can , I want to show you what , what

43:38.360 --> 43:40.600
has happened in the natural image space

43:41.219 --> 43:43.386
and the embedding neighbors , we can ,

43:43.386 --> 43:45.608
particularly for each of the dictionary

43:45.608 --> 43:47.330
elements we can analyze in the

43:47.330 --> 43:49.780
embedding space . What is its neighbors

43:49.790 --> 43:52.012
in that embedding space ? Right . So as

43:52.012 --> 43:54.179
you can see for each of the dictionary

43:54.179 --> 43:56.290
elements , its neighbors are all very

43:56.290 --> 43:59.080
similar uh , uh , uh , you know , this

43:59.090 --> 44:01.090
is a particular stroke , right ? So

44:01.090 --> 44:03.257
that's the , the , the , the neighbors

44:03.257 --> 44:05.423
are all very similar strokes . And the

44:05.423 --> 44:07.757
second one , you know , what is this is ?

44:10.530 --> 44:12.586
It's , um , it's , uh , something in

44:12.586 --> 44:14.586
the letter , uh , in a , in a digit

44:14.586 --> 44:17.639
four , right ? And occasionally it will

44:17.649 --> 44:20.709
be also used in digit seven , right ?

44:20.719 --> 44:23.219
So that , right . So in , in the middle

44:23.229 --> 44:25.229
of the digit seven , right ? So you

44:25.229 --> 44:27.229
have that cross , right ? So in the

44:27.229 --> 44:29.451
embedding space , you can see that even

44:29.451 --> 44:31.562
though they look very differently and

44:31.562 --> 44:33.896
all of them are just this cross , right ?

44:33.896 --> 44:36.062
So that's uh that's pretty shocking uh

44:36.062 --> 44:38.173
original uh moment when I look at the

44:38.173 --> 44:40.285
similarity map , uh you'll find these

44:40.285 --> 44:42.830
neighbors here is from C A , right ? So

44:42.840 --> 44:46.110
C A that that's a great scale of cr and

44:46.120 --> 44:48.287
in this case , you can find that a lot

44:48.287 --> 44:50.509
of the times it , it's , it's basically

44:50.509 --> 44:52.509
the wheel , different angles of the

44:52.509 --> 44:54.676
wheel , they all get placed very close

44:54.676 --> 44:56.842
to each other in the embedding space ,

44:56.842 --> 44:59.009
right ? So in some sense , we start to

44:59.009 --> 45:01.600
discover this parts in the , the

45:01.610 --> 45:04.409
meaning , right ? So here is uh the two

45:04.419 --> 45:06.697
dots . I was like , what , what , what ,

45:06.697 --> 45:08.919
what is that two dots , right ? So it's

45:08.919 --> 45:11.197
uh if you only just look at this patch ,

45:11.197 --> 45:13.419
but when you place all of this neighbor

45:13.419 --> 45:15.586
together , you suddenly realize that's

45:15.586 --> 45:18.679
basically a face of an animal , right ?

45:18.689 --> 45:20.860
So you have this eye and this nose ,

45:20.870 --> 45:23.092
right ? So in the , you can look at the

45:23.092 --> 45:25.314
embedded neighbors that you always have

45:25.314 --> 45:27.370
this , you know , different angles ,

45:27.370 --> 45:29.648
right ? So they are all faces uh right .

45:29.648 --> 45:31.648
So in some sense , you learn a face

45:31.648 --> 45:33.814
cell , right , in that sense . And the

45:33.814 --> 45:36.330
same idea can also be help us explain

45:36.340 --> 45:38.270
what's going on with the complex

45:38.280 --> 45:40.610
pooling in the complex cells , right ?

45:40.620 --> 45:42.898
So I wouldn't get into the detail here .

45:42.898 --> 45:44.898
But um you know , basically you can

45:44.898 --> 45:47.064
explain the why , you know how you can

45:47.064 --> 45:48.969
pull d similar uh a simple cell

45:48.979 --> 45:51.639
together and form of a complex cell and ,

45:51.649 --> 45:53.649
and , and you can show that bipolar

45:53.659 --> 45:55.659
distribution in the , the , the the

45:55.659 --> 45:59.229
response , right . So there um but um I ,

45:59.239 --> 46:02.250
I'll get to the empirical results and

46:02.260 --> 46:04.482
for , for different data sets , right ?

46:04.482 --> 46:07.169
So um m is that these two layer we box

46:07.179 --> 46:09.479
transform . Well , basically , if , if

46:09.489 --> 46:13.310
you use um uh uh uh can , can mean even

46:13.320 --> 46:15.770
uh classifier on top , you get more

46:15.780 --> 46:19.580
than 99% top one accuracy , right ? So

46:19.590 --> 46:21.479
you basically solve that M is the

46:21.479 --> 46:23.701
problem just because that space is very

46:23.701 --> 46:25.812
simple . But again , use this a white

46:25.812 --> 46:28.120
box and was representation learning to

46:28.129 --> 46:30.185
build a two layer transform that can

46:30.185 --> 46:32.518
solve this level . It's actually hasn't ,

46:32.518 --> 46:34.740
hasn't really happened before , right ?

46:34.740 --> 46:36.962
So for C 4 , 10 and C 4 100 you can see

46:36.962 --> 46:39.850
that uh by doing this um 10 years

46:39.860 --> 46:41.804
neighbor again , 10 years neighbor

46:41.804 --> 46:45.489
classifier can get you uh 83 point

46:45.500 --> 46:47.444
something performance , right . So

46:47.449 --> 46:49.782
compared to the state of the art method ,

46:49.782 --> 46:51.782
right , so that the gap is not that

46:51.782 --> 46:53.782
huge . But of course , if you add a

46:53.782 --> 46:56.340
color jittering to the uh to the to the

46:56.350 --> 46:58.350
uh state of the art method , you do

46:58.350 --> 47:00.517
have see a gap . But again , that this

47:00.517 --> 47:02.628
performance is pretty shocking in the

47:02.628 --> 47:04.794
sense that you can actually build this

47:04.794 --> 47:07.659
white box is transform and explain

47:07.669 --> 47:09.725
every step , right . So and then and

47:09.725 --> 47:12.020
build a very competitive uh transform .

47:12.129 --> 47:14.360
I think that's a promising direction

47:14.790 --> 47:17.689
and uh cr 100 is similar . And I um

47:17.699 --> 47:19.532
further , if you compare to this

47:19.532 --> 47:22.090
approach at MT two layers and can

47:22.100 --> 47:24.211
finish in one training epoch . On the

47:24.211 --> 47:26.959
other hand , the uh SSL you know , you

47:26.969 --> 47:29.025
have 18 layers and you train that in

47:29.025 --> 47:32.070
1000 training epochs and you wonder um

47:32.080 --> 47:35.219
can we eventually fully close the gap ,

47:35.229 --> 47:37.340
right ? So to simplify this , so that

47:37.340 --> 47:39.520
potentially make this uh much more

47:39.530 --> 47:41.919
efficient . Yes , we can and I will get

47:41.929 --> 47:43.989
to that point , but let's shift the

47:44.000 --> 47:46.167
gear a little bit to talk about what I

47:46.167 --> 47:48.000
mean by deep and self supervised

47:48.000 --> 47:50.560
learning . The idea is actually the

47:50.570 --> 47:53.030
following . Given the image , we

47:53.040 --> 47:55.699
extract the image patches and then use

47:55.709 --> 47:58.560
color Jitter , right . So uh extract

47:58.570 --> 48:00.879
two image patches and then color Jitter

48:00.889 --> 48:03.000
them and send them into the neural

48:03.010 --> 48:05.479
network in the embedding space . We

48:05.489 --> 48:08.689
want them to be close because they come

48:08.699 --> 48:11.032
from the same image . On the other hand ,

48:11.032 --> 48:12.810
for the image patches come from

48:12.810 --> 48:15.032
different images , we wanted them to be

48:15.032 --> 48:17.255
far apart to be far apart to everything

48:17.255 --> 48:19.421
else , right ? So only the co occurred

48:19.421 --> 48:21.310
ones to be similar , you probably

48:21.310 --> 48:23.421
already start to see the similarity .

48:23.421 --> 48:25.699
Because in in the original formulation ,

48:25.699 --> 48:27.643
Sparse man transform , you put the

48:27.643 --> 48:29.866
cours ones in the embedding space , you

48:29.866 --> 48:31.977
want them to be closed . And you also

48:31.977 --> 48:33.866
want the cour image patches to be

48:33.866 --> 48:36.032
closed except now you are using a deep

48:36.032 --> 48:37.699
neural network and additional

48:37.699 --> 48:40.304
engineering tricks , right . So , and

48:40.314 --> 48:42.425
but there are many such methods . Can

48:42.425 --> 48:45.334
we simplify unify this method and show

48:45.344 --> 48:47.784
what has been learned ? And you know ,

48:47.794 --> 48:50.016
there are many of them and are they all

48:50.016 --> 48:52.127
that different as they seem ? Right .

48:52.127 --> 48:54.127
So it turns out in the first uh aha

48:54.127 --> 48:56.350
moment was that you know , these are uh

48:56.350 --> 48:58.649
method are all motivated by invariants .

48:58.729 --> 49:01.219
You want different image patches from

49:01.229 --> 49:04.080
the image uh from the uh uh from the

49:04.090 --> 49:07.229
same image to be invariant in the

49:07.239 --> 49:09.489
embedding space . But if you continue

49:09.500 --> 49:11.810
that thought , is that OK ? So that

49:11.820 --> 49:15.610
really you don't want to uh uh oh Yeah .

49:15.620 --> 49:17.676
Yeah . Yeah . There's a paper I have

49:17.676 --> 49:19.731
shared in the early announcement I I

49:19.731 --> 49:21.842
think to Kevin . Uh Yeah . Um uh some

49:21.842 --> 49:24.010
of the reference uh to show the major

49:24.020 --> 49:27.830
result here . Um So if you can

49:27.840 --> 49:30.062
get smaller and smaller image patches ,

49:30.062 --> 49:32.284
they cannot be always invariant because

49:32.284 --> 49:33.896
that will give you a trivial

49:33.896 --> 49:36.062
representation , right . So can we get

49:36.062 --> 49:38.118
smaller , smaller , smaller and once

49:38.118 --> 49:40.284
you get smaller and smaller and in the

49:40.284 --> 49:42.284
embedding space to look for all the

49:42.284 --> 49:44.284
neighbors ? And , and here , what I

49:44.284 --> 49:46.396
mean is that for each , for the whole

49:46.396 --> 49:48.451
data set , I extract all these small

49:48.451 --> 49:50.284
image patches , uh uh uh tens of

49:50.284 --> 49:52.562
millions of such image patches , right ?

49:52.562 --> 49:54.451
So and then for each of the image

49:54.451 --> 49:56.173
patches look for the embedding

49:56.173 --> 49:58.118
neighbors in the , in , in this uh

49:58.118 --> 50:01.090
representation phase . And you , when

50:01.100 --> 50:03.156
you get small enough , you will find

50:03.156 --> 50:05.156
that OK , all the neighbors becomes

50:05.156 --> 50:07.211
actually similar parts , right ? You

50:07.211 --> 50:09.300
start to rediscover this uh the part

50:09.310 --> 50:12.620
idea again , right . So and then I can

50:12.629 --> 50:14.796
show you many other examples , right ?

50:14.796 --> 50:17.129
So for each of this image patch , right ,

50:17.129 --> 50:19.129
so you can in in the representation

50:19.129 --> 50:21.240
space find its neighbors , right ? So

50:21.240 --> 50:23.129
they all have the semantic , very

50:23.129 --> 50:25.240
similar semantic meaning . Um And you

50:25.240 --> 50:27.407
can use this idea to even de debug why

50:27.407 --> 50:29.462
the uh the representation give you a

50:29.462 --> 50:31.629
wrong classification ? Because here is

50:31.629 --> 50:33.796
a horse , right . So uh if you look at

50:33.796 --> 50:35.962
this part , it's all from horse , from

50:35.962 --> 50:38.129
this part , you give all to give you a

50:38.129 --> 50:40.429
horse legs with a few cases exception .

50:40.689 --> 50:42.800
Um But if you go here , it's really a

50:42.800 --> 50:45.022
shadow , right ? So shares a lot of the

50:45.022 --> 50:47.290
similarity with the dogs , right . So ,

50:47.300 --> 50:49.750
and here is the image uh when you

50:49.760 --> 50:52.000
classify that uh that's an airplane ,

50:52.010 --> 50:54.177
right ? So and then um but if you look

50:54.177 --> 50:56.232
at the the sky , right ? So a lot of

50:56.232 --> 50:58.288
the the ships will have a lot of the

50:58.288 --> 51:00.510
similar skies . And if you look at this

51:00.510 --> 51:02.677
image patch , right ? So that shares a

51:02.677 --> 51:05.649
very similar salute to a a ship , right ?

51:05.659 --> 51:07.770
So at the end , you know , this , you

51:07.780 --> 51:10.385
have a high chance of uh confusing this

51:10.395 --> 51:12.415
with the ship , I mean , right ? So

51:12.425 --> 51:14.764
that's this part analysis . It's very

51:14.774 --> 51:16.996
informative in that sense to understand

51:16.996 --> 51:19.915
what's going on . Yeah , sorry to

51:19.925 --> 51:22.092
interrupt . Um I think some folks have

51:22.092 --> 51:24.314
a hard stop coming up and our colleague

51:24.314 --> 51:26.258
George Senko , I think has a has a

51:26.258 --> 51:28.203
quick question to ask if that's OK

51:28.203 --> 51:30.314
before he has to go while we have you

51:30.314 --> 51:29.995
on the line , please , please . Yeah .

51:32.209 --> 51:34.590
Yeah . Hi . Ye thank you . Yeah , I

51:34.600 --> 51:36.822
have to stop . I have another important

51:36.822 --> 51:39.659
meeting uh coming up but um maybe

51:39.669 --> 51:41.947
you're gonna cover this . But you know ,

51:41.947 --> 51:44.002
one of the big things that made deep

51:44.002 --> 51:45.613
learning what it is is their

51:45.613 --> 51:49.340
performance on uh uh image net and the

51:49.350 --> 51:52.310
Alex net , right ? They could show and

51:52.320 --> 51:54.209
then , you know , same thing with

51:54.209 --> 51:55.987
transformers and , and language

51:55.987 --> 51:58.300
translation . It was like a noticeable .

51:58.500 --> 52:02.110
So uh what , how are you gonna uh

52:02.120 --> 52:05.310
uh sort of demonstrate that this

52:05.320 --> 52:08.949
impressive body of work is better than

52:08.959 --> 52:10.903
whatever state of the art might be

52:10.903 --> 52:13.015
there ? Or if there's no state of the

52:13.015 --> 52:15.015
art , how are you gonna demonstrate

52:15.015 --> 52:17.790
that it's useful , you know , overall ,

52:18.300 --> 52:20.379
that's awesome question , right ? So

52:20.389 --> 52:22.611
this question is awesome . So I , I was

52:22.611 --> 52:24.667
hoping that someone asked us so it's

52:24.667 --> 52:26.778
actually not better than state of the

52:26.778 --> 52:28.722
art , it's worse , right ? So as I

52:28.722 --> 52:30.945
showed you , there's a gap , right ? So

52:30.945 --> 52:33.167
the the the challenge here , right ? So

52:33.167 --> 52:35.333
the next step I'm really excited about

52:35.333 --> 52:37.949
that is um can we build these white box

52:37.959 --> 52:41.530
model to go back to the Alex net moment ?

52:41.570 --> 52:44.209
Can we build a white box model that can

52:44.219 --> 52:46.929
compete with Alex net , which hasn't

52:46.939 --> 52:48.995
happened before , right . So we have

52:48.995 --> 52:51.050
been using this engineering practice

52:51.050 --> 52:53.217
and you know , it's really working but

52:53.217 --> 52:55.161
uh you know , if we really want to

52:55.161 --> 52:57.161
understand and build the science of

52:57.161 --> 52:59.161
that , how do we de demonstrate the

52:59.161 --> 53:01.383
science ? Right . So and I think the uh

53:01.383 --> 53:05.179
uh one necessary step , it has to be

53:05.189 --> 53:07.389
right . There are only a few people in

53:07.399 --> 53:09.677
the in the few working on this problem ,

53:09.677 --> 53:11.566
but it's absolutely important for

53:11.566 --> 53:13.899
someone to carry the torch . Is that OK ?

53:13.899 --> 53:16.121
Can we for each of this data set one by

53:16.121 --> 53:18.860
one build , we mo we biox model to

53:18.870 --> 53:20.926
compete with the state of the art uh

53:20.926 --> 53:23.300
solutions , right ? Maybe still some

53:23.310 --> 53:25.366
gaps . But let's say for image net ,

53:25.366 --> 53:27.532
can we build a web box model ? Get you

53:27.532 --> 53:29.840
60% top one accuracy ? That's a grand

53:29.850 --> 53:32.017
challenge , right ? So for the science

53:32.017 --> 53:34.100
community , because uh you know , so

53:34.110 --> 53:36.221
far in the past decade , the majority

53:36.221 --> 53:37.889
of the advancement come from

53:37.899 --> 53:40.121
engineering , right ? I I share some of

53:40.121 --> 53:42.288
the engineering background but I think

53:42.288 --> 53:44.343
for this particular pillar , I think

53:44.350 --> 53:46.572
the grand challenge is that we , we are

53:46.572 --> 53:48.600
not at there yet , but um I , I

53:48.610 --> 53:52.139
envision that we should go there if

53:52.149 --> 53:54.399
that makes sense . Yeah . Yeah , I

53:54.409 --> 53:56.679
guess , you know , is there uh a

53:56.689 --> 54:00.439
benchmark or some standard data set

54:00.449 --> 54:03.620
which you might produce that um

54:04.439 --> 54:08.199
uh sort of , you know , uh establishes

54:08.209 --> 54:11.179
some uh something about

54:11.330 --> 54:15.260
representation learning , like what

54:15.270 --> 54:17.610
would be , you know , you could , you

54:17.620 --> 54:19.731
could sort of demonstrate your , your

54:19.731 --> 54:22.520
work achieves a certain performance in

54:22.530 --> 54:24.808
representation learning . I don't know ,

54:24.808 --> 54:27.159
you know , that's uh uh something to

54:27.169 --> 54:29.360
think about . But then , then that's

54:29.370 --> 54:31.370
sort of a concrete thing that , you

54:31.370 --> 54:33.648
know , as opposed to going to Alex net ,

54:33.648 --> 54:36.080
that's an an image net because that's

54:36.090 --> 54:38.419
sort of uh you know , an old problem .

54:39.050 --> 54:40.939
I , I like that you're working on

54:40.939 --> 54:43.272
representation learning what would be a ,

54:43.560 --> 54:45.671
a metric that we could use for that ?

54:46.219 --> 54:49.719
Um I don't fully know , but III I

54:49.729 --> 54:51.840
love your question because uh Euro

54:51.850 --> 54:54.070
simul channel asked exactly the same

54:54.080 --> 54:56.413
question when I was presenting . And he ,

54:56.413 --> 54:58.469
he said , um you know , basically is

54:58.469 --> 55:00.429
that rather than use this standard

55:00.439 --> 55:02.439
machine learning data set , which I

55:02.439 --> 55:04.606
also think is very important , right ?

55:04.606 --> 55:06.828
So because , you know , we , we have to

55:06.828 --> 55:08.883
compete in , in different you know ,

55:08.883 --> 55:08.800
competitive domain , right ? But not

55:08.810 --> 55:10.977
limited in that domain , right ? So is

55:10.977 --> 55:12.977
he asked how , how can you actually

55:12.977 --> 55:14.977
show me the manifold ? Right . So ,

55:14.977 --> 55:17.088
IIII I told him that we are not there

55:17.088 --> 55:19.088
yet , right ? So it's um it's we we

55:19.088 --> 55:20.866
don't stop , still do not fully

55:20.866 --> 55:22.977
understand this space , but we can at

55:22.977 --> 55:25.199
least build some of these um you know ,

55:25.199 --> 55:27.366
in standard machine learning data that

55:27.366 --> 55:26.959
show that competitive performance ,

55:26.969 --> 55:29.399
right ? But again , I , I agree that

55:29.409 --> 55:32.510
potentially would be a data set where

55:32.520 --> 55:34.631
the parts information are out there .

55:34.631 --> 55:36.798
And there's a hierarchical structure .

55:36.798 --> 55:38.798
And we show that this transform can

55:38.798 --> 55:40.742
learn in a hierarchical fashion to

55:40.742 --> 55:42.687
disentangle that all the atoms and

55:42.687 --> 55:44.742
build up the higher structures , you

55:44.742 --> 55:46.853
know , in each of the stage . I think

55:46.853 --> 55:48.964
that's probably the , the some of the

55:48.964 --> 55:50.798
special data set which right now

55:50.798 --> 55:53.229
doesn't exist , right ? So , um yeah ,

55:53.260 --> 55:56.949
so , so I really enjoyed your talk .

55:56.959 --> 55:59.015
I'm gonna take a look at your papers

55:59.015 --> 56:01.379
and uh uh apologize for jumping in ,

56:01.389 --> 56:03.879
but I have to stop right now . All

56:03.889 --> 56:06.000
right . All right , thanks , thanks ,

56:06.000 --> 56:08.280
thanks . Um um All right . Thank you .

56:08.340 --> 56:10.451
Uh Let me continue a little bit . I'm

56:10.451 --> 56:12.562
almost there . I , I should finish uh

56:12.562 --> 56:15.610
very soon . Um Here is that uh OK . So

56:15.620 --> 56:17.731
we have showed that the image patches

56:17.731 --> 56:20.060
can actually show you the , the the

56:20.070 --> 56:23.020
other angle of many vance , right ? So

56:23.030 --> 56:25.689
the variant part in uh deep step was

56:25.699 --> 56:27.905
learning and can we leverage this idea

56:27.915 --> 56:30.084
to build our representation , right .

56:30.094 --> 56:32.316
So here is that for each of the image ,

56:32.316 --> 56:34.538
we extract the image patches and do the

56:34.538 --> 56:36.761
color catering again . But only use the

56:36.761 --> 56:38.872
image patches , not multi scale , not

56:38.872 --> 56:40.705
other large scale you know image

56:40.705 --> 56:42.927
patches but small ones , right . So and

56:42.927 --> 56:44.983
then use the same learning objective

56:44.983 --> 56:47.150
but only use image patches and each of

56:47.150 --> 56:49.364
the image patches is is independently

56:49.375 --> 56:51.375
transformed into the representation

56:51.375 --> 56:53.542
space . And then eventually you do the

56:53.542 --> 56:55.597
aggregation and normalization as the

56:55.597 --> 56:58.550
SMT . And after that , you find that

56:58.560 --> 57:01.449
the patch based training and use this

57:01.459 --> 57:04.179
strategy give you almost exactly the

57:04.189 --> 57:06.399
same perform as the baseline , right .

57:06.409 --> 57:08.576
So use this idea , of course , you can

57:08.576 --> 57:10.798
also improve the baseline in that sense

57:10.798 --> 57:12.576
because you understand that the

57:12.576 --> 57:14.631
underlying representation comes from

57:14.631 --> 57:16.687
image patches and now you can design

57:16.687 --> 57:18.909
better method and significantly improve

57:18.909 --> 57:20.965
the performance . But uh really that

57:20.965 --> 57:23.187
give you an idea , right . So maybe the

57:23.187 --> 57:25.020
representation is uh behind that

57:25.020 --> 57:27.270
success of the self learning , there's

57:27.280 --> 57:29.391
a distributed representation of image

57:29.391 --> 57:31.659
patches behind that , right . So how do

57:31.669 --> 57:33.780
we show that here is a striking uh

57:33.790 --> 57:36.090
result is well , maybe not that

57:36.100 --> 57:38.267
striking some of you probably know the

57:38.267 --> 57:41.679
back uh back net by Mius Bada , right .

57:41.689 --> 57:43.856
So um in supervised learning , right ?

57:43.856 --> 57:45.967
So basically showing that you can use

57:45.967 --> 57:48.250
the image patches to build very strong

57:48.260 --> 57:50.810
supervised representation here . I'm

57:50.820 --> 57:52.876
basically showing the again the same

57:52.876 --> 57:55.098
result in self West learning . It's you

57:55.098 --> 57:58.100
can use 32 by 32 image patches in image

57:58.110 --> 58:00.560
net to build a very strong uh I think

58:00.570 --> 58:04.530
this is a 62% top one accuracy in

58:04.540 --> 58:08.139
image net , use 32 by 32 image patches

58:08.149 --> 58:11.510
right out of that 224 by 224 large

58:11.520 --> 58:13.820
images , right ? So that's pretty uh

58:13.830 --> 58:16.169
surprising that how well this can can

58:16.179 --> 58:19.979
do . Um and further and there are many

58:19.989 --> 58:21.889
such steps tow learning methods

58:21.899 --> 58:25.080
proposed and and um uh they all have

58:25.090 --> 58:27.729
the similar idea flavor to put similar

58:27.739 --> 58:29.906
things uh closer in the representation

58:29.909 --> 58:32.020
space and they work on different anti

58:32.149 --> 58:34.300
collapse mechanisms method . One is

58:34.310 --> 58:36.530
this I I don't , don't go into the

58:36.540 --> 58:38.949
details and that the reference I I will

58:38.959 --> 58:41.850
provide that . Uh but mathematically ,

58:41.860 --> 58:44.879
you can unify these objectives , um you

58:44.889 --> 58:47.010
know under some certain uh certain

58:47.020 --> 58:49.500
conditions and together with several

58:49.510 --> 58:51.949
other works , you can almost fully uni

58:52.139 --> 58:55.419
unify many of such methods mo obyoals

58:55.459 --> 58:58.250
MC M . They , they all share the same

58:58.260 --> 59:00.316
idea . Even sometimes it claims that

59:00.316 --> 59:03.510
the no compressive part uh uh there's

59:03.520 --> 59:05.639
uh implicit con covariance

59:05.649 --> 59:08.060
regularization in the learning dynamics .

59:08.070 --> 59:10.126
You can show that with some analytic

59:10.126 --> 59:12.348
transforms . At the end of the day , it

59:12.348 --> 59:15.229
really show a unified convergence ,

59:15.239 --> 59:17.699
right . So is of these two paradigm in

59:17.820 --> 59:20.550
both of these are basically building a

59:20.560 --> 59:23.229
representation of image patches except

59:23.239 --> 59:25.350
for spars man transform . You use the

59:25.350 --> 59:27.461
sparsity and then the low dimensional

59:27.461 --> 59:30.209
embedding two layers . And on the other

59:30.219 --> 59:32.219
hand , this uh self two we learning

59:32.219 --> 59:34.250
deep self two we learning , you use

59:34.260 --> 59:36.204
these image patches and use a deep

59:36.204 --> 59:38.371
neural network to do that , right . So

59:38.371 --> 59:40.538
and , and they , they , they share the

59:40.538 --> 59:42.810
same mathematical formulation and the

59:42.820 --> 59:45.020
learning of course in smart manifold

59:45.030 --> 59:47.399
transform it , it doesn't need all

59:47.409 --> 59:49.699
those engineering tricks . Uh you can

59:49.709 --> 59:52.959
solve that in one epoch and um uh deep

59:52.969 --> 59:55.120
self learning , you need that 1000

59:55.169 --> 59:57.225
epoch and deep network , but you can

59:57.225 --> 01:00:00.459
close this gap further . By uh uh

01:00:00.469 --> 01:00:03.000
on one hand , the smart man transform ,

01:00:03.010 --> 01:00:04.621
you can build a hierarchical

01:00:04.621 --> 01:00:06.732
representation starting from this low

01:00:06.732 --> 01:00:08.954
level first layer dictionary elements .

01:00:08.954 --> 01:00:11.010
And can you can build on top of that

01:00:11.010 --> 01:00:13.343
and build a hierarchical representation .

01:00:13.343 --> 01:00:15.454
We're not fully quite there yet , but

01:00:15.454 --> 01:00:17.677
there's some evidence and you can learn

01:00:17.677 --> 01:00:19.899
more at a higher level representation ,

01:00:19.899 --> 01:00:22.066
right ? So I need , I think we need to

01:00:22.066 --> 01:00:24.066
work on that more in order to fully

01:00:24.066 --> 01:00:25.843
close the gap . But here's some

01:00:25.843 --> 01:00:28.114
evidence also , once we understand that

01:00:28.125 --> 01:00:30.155
deep self learning performance come

01:00:30.165 --> 01:00:32.276
from these image patients , why don't

01:00:32.276 --> 01:00:34.221
we extract more such image pas per

01:00:34.221 --> 01:00:36.635
image ? Right . So and we can push that

01:00:36.645 --> 01:00:39.114
too extreme and turns out when we do so

01:00:39.125 --> 01:00:42.014
we can make this algorithm converge um

01:00:42.024 --> 01:00:44.584
actually uh within , within five epochs ,

01:00:45.780 --> 01:00:47.947
right . So rather than 1000 epochs and

01:00:47.947 --> 01:00:50.002
you can do that uh you know within ,

01:00:50.002 --> 01:00:52.224
let's say within 10 epochs , right ? So

01:00:52.224 --> 01:00:55.169
it should improve that uh epochs by two

01:00:55.179 --> 01:00:57.179
others magnitude , right ? So I , I

01:00:57.179 --> 01:00:59.235
think that really shows that some of

01:00:59.235 --> 01:01:02.010
these potential uh um uh uh value of

01:01:02.020 --> 01:01:04.409
this uh direction we are not fully

01:01:04.419 --> 01:01:06.909
closing the gap , but uh it can be a

01:01:06.979 --> 01:01:10.330
significant close . Um And finally , I

01:01:10.340 --> 01:01:12.507
would say for , for this , that uh you

01:01:12.507 --> 01:01:15.560
know , the early uh oh yeah . Hey ,

01:01:15.570 --> 01:01:17.570
Jason , hey , oh actually , sorry ,

01:01:17.570 --> 01:01:19.737
maybe I should let you do this slide .

01:01:19.737 --> 01:01:23.060
This is your last one , a couple more ,

01:01:23.070 --> 01:01:25.126
but all of them , they are not , the

01:01:25.126 --> 01:01:26.959
rest are not that engaging , but

01:01:26.959 --> 01:01:30.209
instead it's a playful ones . Um Yeah .

01:01:30.610 --> 01:01:32.721
OK . Well , I guess I just had a kind

01:01:32.721 --> 01:01:34.959
of a , a general question uh which is ,

01:01:34.969 --> 01:01:36.747
you know , the one things we're

01:01:36.747 --> 01:01:38.747
struggling with a lot is that , you

01:01:38.747 --> 01:01:40.747
know , co occurrence . Um It , it's

01:01:40.747 --> 01:01:42.802
great for a lot of things but it's ,

01:01:42.802 --> 01:01:45.610
it's really uh a poor choice uh for

01:01:45.620 --> 01:01:47.600
comparison when you have sort of

01:01:47.610 --> 01:01:49.860
different dimensions of like for

01:01:49.870 --> 01:01:52.037
evaluation that you might want to do .

01:01:52.037 --> 01:01:54.092
And uh so I'm just curious , are you

01:01:54.092 --> 01:01:56.429
thinking about uh like extending your

01:01:56.439 --> 01:01:59.280
ideas to , you know , not just pull

01:01:59.290 --> 01:02:01.512
things that are similar or , you know ,

01:02:01.512 --> 01:02:03.512
or , or co occurring together , but

01:02:03.512 --> 01:02:05.929
sort of maybe pushing things uh you

01:02:05.939 --> 01:02:07.939
know , along one dimension that you

01:02:07.939 --> 01:02:09.883
want to be close together and then

01:02:09.883 --> 01:02:12.050
further away uh that you don't want to

01:02:12.050 --> 01:02:14.161
be together . Uh When , for example ,

01:02:14.161 --> 01:02:16.217
you're trying to , to do things like

01:02:16.217 --> 01:02:18.328
the evaluation . Uh full disclosure .

01:02:18.328 --> 01:02:20.661
I've worked in vision a long time . I'm ,

01:02:20.661 --> 01:02:22.661
I'm an N LP now , but yeah . Yeah .

01:02:22.661 --> 01:02:24.883
Yeah , I think that's awesome idea . Um

01:02:24.883 --> 01:02:27.161
um Yeah , it's , it's hard , it's hard ,

01:02:27.161 --> 01:02:29.328
right ? So what I mean by that is that

01:02:29.330 --> 01:02:31.052
at the beginning , I would say

01:02:31.052 --> 01:02:33.274
representation learning , we aim to , I

01:02:33.274 --> 01:02:35.441
think we aim to build a transform such

01:02:35.441 --> 01:02:37.552
that the structure is explicit . What

01:02:37.552 --> 01:02:39.600
I'm talking about is today . It is

01:02:39.610 --> 01:02:41.780
really about pulling uh reflect the

01:02:41.790 --> 01:02:43.790
similarity . Similarity is just one

01:02:43.790 --> 01:02:46.739
type of , you know , a portion of the

01:02:46.750 --> 01:02:49.649
structure . Uh things can be similar in

01:02:49.659 --> 01:02:51.992
one sense and different from in another ,

01:02:51.992 --> 01:02:54.159
right ? So , and there really are many

01:02:54.159 --> 01:02:56.326
different aspects of the things and we

01:02:56.326 --> 01:02:59.330
uh only this is not enough , right ? So

01:02:59.340 --> 01:03:01.929
what II I today , I just want to show a

01:03:01.939 --> 01:03:04.449
some of the potential of this direction

01:03:04.459 --> 01:03:06.626
is really the what uh you know , fully

01:03:06.626 --> 01:03:10.209
understandable models and um uh and its

01:03:10.219 --> 01:03:12.108
potential in both engineering and

01:03:12.108 --> 01:03:14.163
science . But I , I agree with you .

01:03:14.163 --> 01:03:17.020
It's um um yes . And uh we need to work

01:03:17.030 --> 01:03:19.540
a lot more on that . And fortunate uh

01:03:19.550 --> 01:03:21.494
there are only a few people in the

01:03:21.494 --> 01:03:24.060
whole field working on this direction .

01:03:24.469 --> 01:03:26.979
Um uh I'm one of them carrying the

01:03:26.989 --> 01:03:29.211
torch . Some of the early pioneers like

01:03:29.211 --> 01:03:31.489
Devin Mumford already start to give up .

01:03:31.489 --> 01:03:33.711
I , I remember that I read a po uh blog

01:03:33.711 --> 01:03:35.933
post . You may know Devin and Mumford .

01:03:35.933 --> 01:03:37.711
So who is uh actually the early

01:03:37.711 --> 01:03:39.656
pioneers in patent theory group in

01:03:39.656 --> 01:03:41.878
Brown University , right . So there the

01:03:41.878 --> 01:03:44.045
the the goal is very , very cool and ,

01:03:44.045 --> 01:03:46.000
and um it's building mathematical

01:03:46.010 --> 01:03:48.066
language for all the patterns of the

01:03:48.066 --> 01:03:49.954
signals , right ? So , um I , I'm

01:03:49.954 --> 01:03:52.288
responsible for carrying this direction ,

01:03:52.288 --> 01:03:54.232
but I think there ought to be more

01:03:54.232 --> 01:03:56.232
people working in this direction to

01:03:56.232 --> 01:03:58.399
build the white box models if we can .

01:03:58.399 --> 01:04:00.399
Right . So David , in his blog , he

01:04:00.399 --> 01:04:02.177
said , maybe this generation of

01:04:02.177 --> 01:04:04.343
researchers , we can no longer ask for

01:04:04.343 --> 01:04:06.343
science and you know , maybe in the

01:04:06.343 --> 01:04:08.677
future , we cannot understand our model ,

01:04:08.677 --> 01:04:08.560
right ? Maybe , maybe , right ? So ,

01:04:08.570 --> 01:04:10.810
but uh uh I would say it's too early to

01:04:10.820 --> 01:04:13.360
give up , right ? So Jesus II , I think

01:04:13.370 --> 01:04:16.000
uh we need to work on that more . Uh so

01:04:16.010 --> 01:04:18.199
that potentially build models that can

01:04:18.209 --> 01:04:21.320
compete with transformers , right ? So

01:04:21.370 --> 01:04:23.481
at the moment , that's not the case ,

01:04:23.481 --> 01:04:25.092
right ? So we are still even

01:04:25.092 --> 01:04:26.981
convolution in your networks . We

01:04:26.981 --> 01:04:28.981
there's still gaps , right ? So for

01:04:28.981 --> 01:04:28.739
transformers , I think the gaps would

01:04:28.750 --> 01:04:31.860
be even larger , but that's a very

01:04:31.870 --> 01:04:34.780
promising direction . And though it's

01:04:34.790 --> 01:04:37.379
one , it's extremely hard compared to

01:04:37.389 --> 01:04:39.389
deep learning papers , a lot of the

01:04:39.389 --> 01:04:41.445
times you can publish couple of deep

01:04:41.445 --> 01:04:43.333
learning papers , but the web box

01:04:43.333 --> 01:04:45.222
models try to develop some of the

01:04:45.222 --> 01:04:47.500
understanding and why that takes a lot ,

01:04:47.500 --> 01:04:49.729
a lot longer , uh usually , right ? So

01:04:49.739 --> 01:04:52.310
then I I joke that it's P in a

01:04:52.320 --> 01:04:55.580
different clock I appreciate your

01:04:55.590 --> 01:04:58.770
steadfast determination to uh to tackle

01:04:58.780 --> 01:05:02.320
a very hard problem . I will complement

01:05:02.330 --> 01:05:05.699
that with um build world models , right ?

01:05:05.709 --> 01:05:08.669
So which is purely engineering . And at

01:05:08.679 --> 01:05:10.957
the end I will talk about that , right .

01:05:10.957 --> 01:05:13.389
So , uh OK , so for this one , you know ,

01:05:13.399 --> 01:05:15.288
in the early stage , we have some

01:05:15.288 --> 01:05:17.510
insights about the filters , right ? So

01:05:17.520 --> 01:05:19.576
in computational neuroscience and in

01:05:19.576 --> 01:05:21.576
the high level , we have some ideas

01:05:21.576 --> 01:05:23.798
about , you know , there should be some

01:05:23.798 --> 01:05:25.964
concepts like objects and phases start

01:05:25.964 --> 01:05:27.687
to emerge . And we used to not

01:05:27.687 --> 01:05:29.798
understand how this uh is uh formed .

01:05:29.798 --> 01:05:31.964
And this spark manifold transform give

01:05:31.964 --> 01:05:34.131
you an idea how you can hierarchically

01:05:34.131 --> 01:05:36.250
build a transform that disentangle

01:05:36.260 --> 01:05:38.427
these simple elements and build on top

01:05:38.427 --> 01:05:40.149
of that and build more complex

01:05:40.149 --> 01:05:42.439
structures out of it in a fully

01:05:42.449 --> 01:05:44.560
understandable fashion . Right ? So ,

01:05:44.560 --> 01:05:47.000
and uh some of the uh my uh these works

01:05:47.010 --> 01:05:49.939
have been verified supported by um uh

01:05:49.949 --> 01:05:52.500
compu uh the the from the uh

01:05:52.510 --> 01:05:55.250
neuroscience uh community as well

01:05:55.260 --> 01:05:57.482
recent , right , showing that the brain

01:05:57.489 --> 01:05:59.879
is leveraging very similar strategy ,

01:05:59.889 --> 01:06:03.159
right . So , um to achieve that um all

01:06:03.169 --> 01:06:05.391
right , to summarize the main points of

01:06:05.391 --> 01:06:07.558
today's talk , I apologize . I've been

01:06:07.558 --> 01:06:09.780
slightly over time . And uh uh the main

01:06:09.780 --> 01:06:12.050
point is that uh as representation

01:06:12.110 --> 01:06:14.110
transform , we can derive that from

01:06:14.110 --> 01:06:16.166
neural and Staal principle and build

01:06:16.166 --> 01:06:18.709
competitive uh uh uh algorithms uh

01:06:18.719 --> 01:06:21.449
better than we thought , right . So ,

01:06:21.459 --> 01:06:24.260
and when I DM is uh uh from sparse

01:06:24.270 --> 01:06:26.214
coding can tell the data space and

01:06:26.214 --> 01:06:27.937
provide a support and spectral

01:06:27.937 --> 01:06:30.103
embedding can establish the similarity

01:06:30.103 --> 01:06:32.159
on top of that uh support . On the

01:06:32.169 --> 01:06:34.058
other hand , we can simply find a

01:06:34.058 --> 01:06:36.870
unified deep and learning and uh build

01:06:36.879 --> 01:06:39.280
a a and show that our distributed

01:06:39.290 --> 01:06:41.239
representation of image patches is

01:06:41.250 --> 01:06:43.219
behind that . And leveraging these

01:06:43.229 --> 01:06:46.270
ideas , we can unify many such methods

01:06:46.300 --> 01:06:48.810
and also improve them significantly .

01:06:48.820 --> 01:06:51.850
Right . So , and these two paradigm ,

01:06:51.899 --> 01:06:53.929
they share surprisingly the same

01:06:53.939 --> 01:06:56.209
mathematical objective and their

01:06:56.219 --> 01:06:58.050
engineering gap can also be

01:06:58.060 --> 01:07:01.919
significantly closed . Um And I

01:07:01.929 --> 01:07:03.929
want to talk a little bit about the

01:07:03.929 --> 01:07:05.707
engineering part and the future

01:07:05.707 --> 01:07:07.929
direction . And on one hand , I think I

01:07:07.929 --> 01:07:10.096
have talked enough about uh we need to

01:07:10.096 --> 01:07:12.096
close these gaps for the building a

01:07:12.096 --> 01:07:14.318
white box hierarchical representation ,

01:07:14.318 --> 01:07:16.651
right ? So I wouldn't stress that again .

01:07:16.651 --> 01:07:18.707
I think the grand challenge for this

01:07:18.707 --> 01:07:20.707
generation of science researcher is

01:07:20.707 --> 01:07:23.149
that can we leverage our understanding ,

01:07:23.159 --> 01:07:26.750
build something that is um white box

01:07:26.760 --> 01:07:28.927
and yet can compete with deep learning

01:07:28.927 --> 01:07:31.093
and algorithm because we don't want to

01:07:31.093 --> 01:07:32.760
see to say OK , this is fully

01:07:32.760 --> 01:07:34.840
understandable , you know , but it

01:07:34.850 --> 01:07:36.739
doesn't perform as well as a deep

01:07:36.739 --> 01:07:38.850
learning models , right ? So and then

01:07:38.850 --> 01:07:40.961
uh which one would I choose ? I would

01:07:40.961 --> 01:07:43.183
probably choose deep learning if I want

01:07:43.183 --> 01:07:45.183
to work on any practical problems ,

01:07:45.183 --> 01:07:47.517
right ? So science , science is science ,

01:07:47.517 --> 01:07:49.739
but uh when it comes to engineering , I

01:07:49.739 --> 01:07:51.683
think we need to have an objective

01:07:51.683 --> 01:07:53.739
evaluation , right ? So whatever the

01:07:53.739 --> 01:07:55.906
works better and , you know , that's a

01:07:55.906 --> 01:07:58.072
joke that Jeff Jeff Hinton made . II ,

01:07:58.072 --> 01:07:59.850
I think I would probably prefer

01:07:59.850 --> 01:08:01.961
anything that , uh , perform better .

01:08:01.961 --> 01:08:03.628
Right . Even that's not fully

01:08:03.628 --> 01:08:05.628
understandable if it works better .

01:08:05.628 --> 01:08:07.850
Right . So , why should I shouldn't I ,

01:08:07.850 --> 01:08:10.017
um , yeah , but , uh , you know , in ,

01:08:10.017 --> 01:08:12.183
that's been said , I , I think we need

01:08:12.183 --> 01:08:14.183
to work on that web box model . But

01:08:14.183 --> 01:08:16.128
further , I think I want to talk a

01:08:16.128 --> 01:08:18.072
little bit about the the the other

01:08:18.072 --> 01:08:20.509
direction which is actually the world

01:08:20.520 --> 01:08:22.500
models , right ? So it's the other

01:08:22.509 --> 01:08:25.040
pillars of my lab and it's it's

01:08:25.049 --> 01:08:27.290
building the world models that this is

01:08:27.299 --> 01:08:29.660
at the engineering front here , right ?

01:08:29.669 --> 01:08:31.899
So how do we build models that can go

01:08:31.910 --> 01:08:33.966
beyond perception ? Because a lot of

01:08:33.966 --> 01:08:36.077
those type things I have talked today

01:08:36.077 --> 01:08:38.709
were about perception and there's also

01:08:38.720 --> 01:08:41.500
motor memory or long term prediction

01:08:41.509 --> 01:08:43.676
explorations , right ? So at the end ,

01:08:43.676 --> 01:08:45.731
it's really about how do we actually

01:08:45.731 --> 01:08:47.898
build this engineering system that can

01:08:47.898 --> 01:08:49.731
understand the world with common

01:08:49.731 --> 01:08:52.659
knowledge here . I want to specifically

01:08:52.668 --> 01:08:55.199
zoom into the the the motor part

01:08:55.208 --> 01:08:57.729
sensory is basically given a entangled

01:08:57.738 --> 01:08:59.969
signal and go into the representation

01:08:59.979 --> 01:09:02.257
space and bring the structure explicit .

01:09:02.257 --> 01:09:04.090
On the other hand , the motor is

01:09:04.090 --> 01:09:06.146
started with some simple command and

01:09:06.146 --> 01:09:08.368
then translate that to very complicated

01:09:08.429 --> 01:09:11.258
uh motor actions , embodied actions ,

01:09:11.488 --> 01:09:13.579
right ? So can we also hope for a

01:09:13.588 --> 01:09:16.008
representation space in that regime ?

01:09:16.500 --> 01:09:18.500
And there's some early evidence , I

01:09:18.500 --> 01:09:20.611
think that's that's actually the case

01:09:20.611 --> 01:09:23.000
here you have the a a agent , a very

01:09:23.009 --> 01:09:25.176
simple half Cheetah , right ? So , and

01:09:25.176 --> 01:09:28.390
then uh doing um uh uh we want to embed

01:09:28.399 --> 01:09:30.621
all these different action in the laten

01:09:30.621 --> 01:09:33.029
space so that you can possibly reuse

01:09:33.040 --> 01:09:35.879
this policy network for many other

01:09:35.890 --> 01:09:37.501
different tasks and show the

01:09:37.501 --> 01:09:39.890
generalization , right ? So , uh um

01:09:39.899 --> 01:09:42.390
here for each of the task , you want to

01:09:42.399 --> 01:09:45.270
learn a different latent embedding and

01:09:45.279 --> 01:09:47.640
you , you train many such uh uh tasks .

01:09:48.339 --> 01:09:50.759
Uh you can run , run faster and faster

01:09:50.770 --> 01:09:52.937
and faster until the patterns start to

01:09:52.937 --> 01:09:54.770
change , right ? So , you know ,

01:09:54.770 --> 01:09:56.659
walking slowly and walk , running

01:09:56.659 --> 01:09:58.720
faster , it's really not the same

01:09:58.729 --> 01:10:00.951
action , right ? So the pattern is kind

01:10:00.951 --> 01:10:03.062
of different and further , you can do

01:10:03.062 --> 01:10:05.340
different uh strange behaviors , right ?

01:10:05.340 --> 01:10:07.285
So , and uh train your agent to do

01:10:07.285 --> 01:10:09.618
different things in the embedding space .

01:10:09.618 --> 01:10:11.970
Um And then once we learn this gene

01:10:11.979 --> 01:10:15.089
multitasks in it have a latent space .

01:10:15.100 --> 01:10:17.959
And the question is , can we ask the

01:10:17.970 --> 01:10:20.303
generalization we see in the perception ,

01:10:20.303 --> 01:10:22.526
can we also hope for the generalization

01:10:22.526 --> 01:10:24.748
in the action space ? And the answer is

01:10:24.748 --> 01:10:27.970
yes . In uh uh here you have a running

01:10:27.979 --> 01:10:31.500
half Cheetah , 1 m per 2nd and 2 m per

01:10:31.509 --> 01:10:33.620
second . These are the tasks that has

01:10:33.620 --> 01:10:35.787
been specifically trained on , right ?

01:10:35.787 --> 01:10:38.200
So now you want to have a very simple

01:10:38.209 --> 01:10:40.270
generalization which has never been

01:10:40.279 --> 01:10:43.180
trained on is 1.5 m per second , right ?

01:10:43.189 --> 01:10:45.245
So , and then you can simply just do

01:10:45.245 --> 01:10:47.467
that in the latent space , do an interp

01:10:47.649 --> 01:10:49.870
interpolation in that latent space and

01:10:49.879 --> 01:10:52.046
go through your policy network and you

01:10:52.046 --> 01:10:54.350
suddenly find a 1.5 m per second ,

01:10:54.359 --> 01:10:56.526
which is not that surprising , right ?

01:10:56.526 --> 01:10:59.310
So maybe it should work , right ? So uh

01:10:59.319 --> 01:11:01.850
but what about that ? Right ? So you

01:11:01.859 --> 01:11:05.149
have to walk and stand , jump and run .

01:11:05.390 --> 01:11:07.612
What if you , you interpret them in the

01:11:07.612 --> 01:11:09.799
representation space and see what

01:11:09.810 --> 01:11:11.921
happens in the uh when you go through

01:11:11.921 --> 01:11:13.866
the policy , that's what happens ,

01:11:13.866 --> 01:11:16.032
right ? So you , you have walk and you

01:11:16.032 --> 01:11:18.143
have a stand and then you you do this

01:11:18.143 --> 01:11:20.310
walk stand , right ? So you have never

01:11:20.310 --> 01:11:22.421
been trained on that , right ? So God

01:11:22.421 --> 01:11:24.699
knows whether this should work , right ?

01:11:24.699 --> 01:11:26.810
So , but when you combine that in the

01:11:26.810 --> 01:11:29.089
representative , it worked so very ,

01:11:29.100 --> 01:11:32.100
very simple and early stage um evidence .

01:11:32.149 --> 01:11:34.439
But I think that in the action space ,

01:11:34.450 --> 01:11:36.629
we can also hope for uh some

01:11:36.640 --> 01:11:39.569
generalization and representation . If

01:11:39.580 --> 01:11:41.802
that's also the case , we can now start

01:11:41.802 --> 01:11:44.024
to imagine that we close the loop , the

01:11:44.024 --> 01:11:46.209
representation space of sensory and

01:11:46.220 --> 01:11:48.442
representations of the motor . And then

01:11:48.442 --> 01:11:50.553
on top of it , we build that memory ,

01:11:50.553 --> 01:11:52.720
you know our chat GP T , right ? So it

01:11:52.729 --> 01:11:55.759
will deal with this uh perception and

01:11:55.770 --> 01:11:57.992
action and build that prediction of the

01:11:57.992 --> 01:11:59.937
future . Of course , I didn't talk

01:11:59.937 --> 01:12:01.881
about the exploration , right ? So

01:12:01.881 --> 01:12:03.937
that's where you have to how , how ,

01:12:03.937 --> 01:12:06.159
what's the strategy , what's the reward

01:12:06.159 --> 01:12:08.214
that you will , what's the intrinsic

01:12:08.214 --> 01:12:10.214
reward ? How you , you explore your

01:12:10.214 --> 01:12:12.370
space , maybe curiosity avoid some

01:12:12.379 --> 01:12:14.790
dangers , you know , all these , right ?

01:12:14.799 --> 01:12:17.500
So that's totally another topic . But I

01:12:17.509 --> 01:12:19.879
think there could be a unified future

01:12:19.890 --> 01:12:22.112
that we can build the world models with

01:12:22.112 --> 01:12:24.500
these latent representation so that you

01:12:24.509 --> 01:12:27.660
can have a very generic agent that will

01:12:27.669 --> 01:12:30.080
operate in the real world . Right ? So ,

01:12:30.089 --> 01:12:31.700
and accomplish a more uh and

01:12:31.700 --> 01:12:35.180
generalized to new tasks uh quickly .

01:12:36.080 --> 01:12:38.302
So yeah , that's uh at the end , I want

01:12:38.302 --> 01:12:41.580
to echo uh Lean's uh recent white paper ,

01:12:41.859 --> 01:12:44.081
uh word model , right ? So I believe in

01:12:44.081 --> 01:12:46.192
that direction , right ? So that's um

01:12:46.192 --> 01:12:48.979
my group would try to inherit it from

01:12:48.990 --> 01:12:51.549
Bruno and try to find this principles .

01:12:51.649 --> 01:12:53.816
But also compliment that with uh Yen's

01:12:53.816 --> 01:12:56.149
perspective , build engineering , right ?

01:12:56.149 --> 01:12:58.260
Engineering goes first , right ? So ,

01:12:58.260 --> 01:13:00.427
and potentially these pillars can even

01:13:00.427 --> 01:13:02.760
help each other to very extreme , right ?

01:13:02.760 --> 01:13:04.538
The frontier of engineering and

01:13:04.538 --> 01:13:06.427
frontier of science in South West

01:13:06.427 --> 01:13:08.740
learning . I finished my talk and any

01:13:08.750 --> 01:13:11.339
further question , uh apologize for the

01:13:11.350 --> 01:13:15.180
being a slightly over time . Oh ,

01:13:15.189 --> 01:13:17.000
thank you . Thank you . Wow .

01:13:20.830 --> 01:13:24.350
Um On the uh on the walk stand , uh I

01:13:24.359 --> 01:13:26.581
noticed you had sort of the hind leg um

01:13:26.581 --> 01:13:29.109
was doing the , the walk , the motion

01:13:29.120 --> 01:13:33.040
from walk . Um Yeah . Uh

01:13:33.049 --> 01:13:34.882
assuming that isn't , that isn't

01:13:34.882 --> 01:13:36.882
something that's just necessary for

01:13:36.882 --> 01:13:38.827
balance and that it is a sort of a

01:13:38.827 --> 01:13:40.938
vestigial uh kind of thing . Is there

01:13:40.938 --> 01:13:43.640
any sort of a procedure ? Um is there

01:13:43.649 --> 01:13:45.593
any sort of procedures for sort of

01:13:45.593 --> 01:13:48.479
identifying and pruning vestigial

01:13:48.490 --> 01:13:50.712
actions that are no longer necessary at

01:13:50.712 --> 01:13:52.268
the interpolated uh for the

01:13:52.268 --> 01:13:54.434
interpolated task such that they don't

01:13:54.434 --> 01:13:58.399
just propagate down the uh I like your

01:13:58.410 --> 01:14:00.600
comments . I think you have good . I

01:14:00.609 --> 01:14:04.000
didn't realize that uh myself , right ?

01:14:04.009 --> 01:14:06.580
So just to repeat that you can see that

01:14:06.589 --> 01:14:08.811
this is almost like a walking , right ?

01:14:08.919 --> 01:14:11.270
So , but in uh at the same time , it's

01:14:11.279 --> 01:14:13.446
the balancing , right ? So if you look

01:14:13.446 --> 01:14:15.557
at this leg , it's almost like a walk

01:14:15.557 --> 01:14:17.668
compared to here , it's relatively ST

01:14:17.668 --> 01:14:20.950
um I , I don't know . My gut feeling is

01:14:20.959 --> 01:14:23.299
that regularization can certainly help ,

01:14:23.310 --> 01:14:25.421
right ? So because reg regularization

01:14:25.421 --> 01:14:27.643
is that when we do not have enough data

01:14:27.643 --> 01:14:29.810
regularization can a lot of times help

01:14:29.810 --> 01:14:31.421
the generalization . Another

01:14:31.421 --> 01:14:33.643
complementary perspective is that maybe

01:14:33.643 --> 01:14:35.866
we should just scale , right ? So scale

01:14:35.866 --> 01:14:37.810
this approach and so that you , we

01:14:37.810 --> 01:14:40.669
train uh more tasks , more self-support

01:14:40.859 --> 01:14:43.799
wise task , maybe hundreds of thousands

01:14:43.810 --> 01:14:45.588
such tasks . And then maybe the

01:14:45.588 --> 01:14:47.600
generalization would also improve .

01:14:48.120 --> 01:14:50.120
That's my , my feeling , right ? So

01:14:50.120 --> 01:14:52.319
regularization is that we do not have

01:14:52.330 --> 01:14:54.919
enough knowledge , but if we echo rich

01:14:54.930 --> 01:14:57.097
sentence uh bitter lesson , right ? So

01:14:57.097 --> 01:15:00.080
we should always scale . Uh uh um I , I

01:15:00.089 --> 01:15:02.200
think both of this idea could be , be

01:15:02.209 --> 01:15:03.479
useful . Yeah .

01:15:12.410 --> 01:15:13.549
Hi . Um

01:15:17.910 --> 01:15:20.021
Other other questions , folks , while

01:15:20.021 --> 01:15:22.680
we have uh you be on the line , Doctor

01:15:22.689 --> 01:15:23.209
Chen

01:15:27.189 --> 01:15:29.356
Scott , I don't know if you wanna , if

01:15:29.356 --> 01:15:31.522
you wanna say anything since you , you

01:15:31.522 --> 01:15:33.356
introduced us or close us out or

01:15:33.356 --> 01:15:35.467
anything . So , I it's kind of you to

01:15:35.467 --> 01:15:37.689
say that I introduced us . I uh uh my ,

01:15:37.720 --> 01:15:39.831
my internet connection stutter . So I

01:15:39.831 --> 01:15:41.887
apologize for that . No , no , III I

01:15:41.887 --> 01:15:44.109
just appreciate , thanks again , you be

01:15:44.109 --> 01:15:46.220
for , for calling in and joining us .

01:15:46.220 --> 01:15:46.209
This is great . I , I know a bunch of

01:15:46.220 --> 01:15:48.387
folks in , the , bunch of folks in the

01:15:48.387 --> 01:15:50.720
team are gonna read your papers now and ,

01:15:50.720 --> 01:15:52.887
and I , I , I'm gonna be one of them .

01:15:52.887 --> 01:15:52.799
OK . Thank you . Thank you . I'll be

01:15:52.810 --> 01:15:54.939
happy to uh chat more about these ,

01:15:54.950 --> 01:15:57.061
right ? So I think a lot of these are

01:15:57.061 --> 01:15:59.283
still we need more effort on that , you

01:15:59.283 --> 01:16:01.228
know , this is still very simple ,

01:16:01.228 --> 01:16:03.394
right ? So my dream is that eventually

01:16:03.394 --> 01:16:05.506
on one hand , we can have these fully

01:16:05.506 --> 01:16:07.672
white box models that are , you know ,

01:16:07.672 --> 01:16:09.672
fully understandable . On the other

01:16:09.672 --> 01:16:11.728
hand , let's scale this , you know ,

01:16:11.728 --> 01:16:13.506
for the , the the um real world

01:16:13.506 --> 01:16:15.228
robotics , right ? So that's a

01:16:15.228 --> 01:16:17.394
potentially a direction so that we can

01:16:17.394 --> 01:16:19.561
have a very general agent , right ? So

01:16:19.561 --> 01:16:21.672
these are the two pillars , I hope to

01:16:21.672 --> 01:16:23.672
establish more uh with my new lab ,

01:16:23.672 --> 01:16:25.783
right ? So I used to be just myself ,

01:16:25.783 --> 01:16:27.894
right ? So , uh and my club readers ,

01:16:27.894 --> 01:16:30.061
but now I , I start to establish a lab

01:16:30.061 --> 01:16:32.283
and start to hi hire students , right ?

01:16:32.283 --> 01:16:34.061
So potentially we can push this

01:16:34.064 --> 01:16:36.535
direction further . Yeah , together .

01:16:36.544 --> 01:16:38.766
Yeah . Yeah , that's , that's great . I

01:16:38.766 --> 01:16:40.822
know . I know uh uh if Doug's on the

01:16:40.822 --> 01:16:43.100
line , he likes to , to talk about how ,

01:16:43.100 --> 01:16:45.377
you know , you , you spend time on the ,

01:16:45.377 --> 01:16:44.875
at the blackboard in the morning and

01:16:44.884 --> 01:16:46.662
then you build something in the

01:16:46.662 --> 01:16:48.828
afternoon . And so that's the , that's

01:16:48.828 --> 01:16:50.551
the two pillars that the way I

01:16:50.551 --> 01:16:50.015
understand what you're saying . So I

01:16:50.024 --> 01:16:52.415
think it , it sounds like a great way

01:16:52.424 --> 01:16:55.290
forward to me . OK . Yeah . Yeah ,

01:16:55.299 --> 01:16:57.299
thank you . Right . So I , I I'd be

01:16:57.310 --> 01:17:00.540
happy to uh chat more . It was a

01:17:00.549 --> 01:17:03.549
wonderful talk . Thank you . Oh , thank

01:17:03.560 --> 01:17:04.560
you . Thank you .

01:17:10.919 --> 01:17:13.750
All right . So , all right then . No

01:17:13.859 --> 01:17:16.600
more . Yeah , let's schedule other

01:17:16.609 --> 01:17:18.609
meetings , right ? So if we want to

01:17:18.609 --> 01:17:20.831
brainstorm some of the further research

01:17:20.831 --> 01:17:23.053
ideas in the future , right ? So we can

01:17:23.053 --> 01:17:25.109
do that . Um But it's really a great

01:17:25.109 --> 01:17:27.276
pleasure to be here today and meet all

01:17:27.276 --> 01:17:29.700
of you , a broader audience . Um And

01:17:29.709 --> 01:17:31.653
thanks for having me here . Well ,

01:17:31.870 --> 01:17:34.203
thanks again for joining us . All right .

01:17:34.203 --> 01:17:36.537
Hope you all have a good weekend . Yeah .

01:17:36.537 --> 01:17:38.759
Thank you . I'm going to teach my first

01:17:38.759 --> 01:17:40.814
class and my , my good time will end

01:17:40.814 --> 01:17:42.870
soon . Oh , well , best of luck with

01:17:42.870 --> 01:17:46.870
that . Ok . Ok . Thank you . See

01:17:46.879 --> 01:17:47.910
you guys . Bye bye .

